Using multi-backend model access as the setting, this article explains how an AI Gateway handles resource scheduling, capacity control, failure switching, distributed state, usage accounting, and enterprise cost governance.
Short
16 min read
Part of the column “Technical Systems” · Chapter 9
Using a multi-layer execution service as an example, this article breaks down the multiple levels of locality across request identity, edge routing, execution resources, and backend caches, then presents reusable approaches to routing, invalidation, failover, and observability.
Short
13 min read
Part of the column “Technical Systems” · Chapter 7
Starting from the fault structure of a multi-node execution system, this article reviews how requests, state, storage, performance, and failures gradually cross boundaries, and discusses how the next architecture should be reorganized.
Short
19 min read
Part of the column “Technical Systems” · Chapter 6
Starting from the different questions asked by analysis, operations, and engineering, this article organizes how a data center for a multi-source business platform should divide Tabs, define metrics, structure read and write paths, and handle billing, reconciliation, growth, and cost.
Starting with referral rewards and transaction systems, this article designs reusable, auditable marketing infrastructure across campaigns, audiences, entitlements, budgets, attribution, risk, experimentation, messaging, and measurement.
Starting from a multi-system integration experience, this essay distinguishes creation, task efficiency, organizational productivity, and business value—and asks where value and cost actually come from when AI enters a complex system.
Short
5 min read
Part of the column “Engineering & AI Judgment” · Chapter 10
Starting from the observability problems of a small API service, this article distills the minimum implementation of request IDs, distributed Traces, service discovery, health monitoring, log queries, reliability, and infrastructure governance.
Based on long-task use, this paper examines how an Agent can adapt state, evidence, and action feedback to uncertainty, waiting, failure, and takeover needs, while documenting how delegated execution changes human pressure, trust, and capability boundaries.
Based on ChatLab design work, this paper records how presence, response, waiting, and delivery patterns from human collaboration can make Agent runtime states easier to understand; it is an HCI design observation about feedback, evidence, intervention, and result confirmation.
This paper examines task decomposition, subtask capability classification, model selection, and bounded execution-time fallback, while reducing the broader Agent scheduling problem to an engineering slice of dsh-quota-router.
A policy-only DeepSeek Harness plugin that builds ordered, explainable, auditable model candidate chains per task and advances along them within bounded failure rules.
A publicly reusable metric dictionary: from requests and users to tasks, explaining how availability, error, latency, performance, and feedback data should be defined, combined, and interpreted.
Short
13 min read
Part of the column “Data Metrics Guide” · Chapter 2