Listen to this article
Narrated by Charlotte · The Noble House
The hum of a GPU cluster is a dynamic conversation of heat and electricity. In the early days of large language model infrastructure, this conversation was chaotic. Systems relied on decentralized, feedback-driven mechanisms. Load balancers smoothed engine-reported signals into scores and nudged routing weights to balance constrained engines [3]bestblogs.devRouting LLM Inference in Production: From Engine Signals to PolicyOpenAI inference engineers Lu Zhang and Qianru Lao recount how IRB evolved from a feedback-loop load balancer into a control-plane and data-plane design whose global optimizer assigns routing weights to minimize end-to-end latency under…Open source ↗. These systems functioned as reactive reflexes, adjusting to immediate conditions without a holistic view of the fleet. As the scale of AI deployments exploded, the limitations of this emergent behavior became critical. The industry’s transition from these legacy systems to centralized control-plane architectures marks a fundamental shift in how artificial intelligence is governed, moving from implicit, reactive adjustments to explicit, policy-driven optimization [9]agihunt.infoInside OpenAI's inference routing: why… · AGI HuntThe replacement is a control plane + data plane design: a control plane with a global view of every CPU cluster and GPU engine, and a per-cluster data plane that answers the one synchronous question—which engine serves this request—from a…Open source ↗.
This evolution is best illustrated by the transformation of OpenAI’s Inference Routing Balancer (IRB). Lu Zhang and Qianru Lao, inference engineers at OpenAI, have detailed how the system evolved from a feedback-loop load balancer into a control-plane and data-plane design [3]bestblogs.devRouting LLM Inference in Production: From Engine Signals to PolicyOpenAI inference engineers Lu Zhang and Qianru Lao recount how IRB evolved from a feedback-loop load balancer into a control-plane and data-plane design whose global optimizer assigns routing weights to minimize end-to-end latency under…Open source ↗. This architectural change allows for a global optimizer that assigns routing weights to minimize end-to-end latency under capacity, health, and cache constraints [9]agihunt.infoInside OpenAI's inference routing: why… · AGI HuntThe replacement is a control plane + data plane design: a control plane with a global view of every CPU cluster and GPU engine, and a per-cluster data plane that answers the one synchronous question—which engine serves this request—from a…Open source ↗. The move from local proportional control to global optimization reflects a broader industry trend toward treating AI infrastructure with the same rigor as financial or logistics networks, where predictability and efficiency are paramount. This transition underscores the necessity of auditing the optimizer's objective functions and verifying alignment with human-intended utility and safety constraints [1]ai-redteam.comRouting LLM Inference in Production: From Engine Signals to Policy — Qianru Lao & Lu Zhang, OpenAIThe routing weights inside OpenAI's inference load balancer used to come out of a feedback loop. Engines reported signals, a controller smoothed them into a score, compared it to the fleet average, and nudged each weight up or down.Open source ↗.
The Legacy of Feedback-Loop Balancing
The original design of inference routing relied on a decentralized feedback loop. In this model, engines reported signals regarding their current state, such as load or latency. A controller then smoothed these signals into a score, compared them to the fleet average, and nudged each weight up or down. This approach was reactive and local, meaning that decisions were made based on immediate, observed conditions rather than a holistic view of the system. While effective for smaller clusters, this method struggled with the complexity of large-scale deployments where interdependencies between CPU clusters and GPU engines created unpredictable bottlenecks.
The reliance on continuous weight nudging introduced instability. As the system attempted to balance load, it often overcorrected, leading to oscillations in performance. Furthermore, the lack of a global view meant that the system could not account for cache states or health metrics across the entire fleet. This limitation became increasingly problematic as the demand for low-latency responses grew. The legacy system’s inability to anticipate traffic patterns or optimize for end-to-end latency rather than just local load balancing highlighted the need for a more sophisticated architecture. The practical agent stack of routing, usable context, and security built in requires a foundation that can handle these complexities, which the feedback-loop model could not provide [7]qabash.comRouting LLM Inference in Production: From Engine Signals to Policy ...The practical agent stack is routing, usable context, and security built in. The routing weights inside OpenAI's inference load balancer used to come out of a feedback loop.Open source ↗.
The transition away from this model was not immediate but gradual, driven by the increasing complexity of the workloads. Engineers observed that the feedback loop often failed to converge on an optimal state, especially during peak traffic periods. The system would spend more time correcting its own errors than serving requests efficiently. This inefficiency was not just a technical issue but a governance one, as the emergent behavior of the system was difficult to audit or explain. The lack of explicit policy meant that the system’s decisions were opaque, making it hard to ensure that routing decisions aligned with safety or cost constraints. The video 'Routing LLM Inference in Production: From Engine Signals to Policy' was published by AI Red Team on September 19, 2026, capturing this pivotal moment in infrastructure history [6]ai-redteam.comRouting LLM Inference in Production: From Engine… | AI Red TeamMy takeaway: Routing LLM Inference in Production: From Engine Signals to Policy — Qianru Lao & Lu Zhang, OpenAI is a governance signal. The practical read is to map the policy language into controls, audit evidence, ownership, and…Open source ↗.
Compass Predictive Analytics
Compass Predictive Analytics

Control-Plane and Data-Plane Architecture
The replacement for the legacy system is a control-plane and data-plane design. This architecture separates the decision-making process from the execution process. The control plane maintains a global view of every CPU cluster and GPU engine, allowing it to make informed decisions about routing. The data plane, located in each cluster, answers the synchronous question of which engine serves a request from a cached snapshot of routing weights. This separation allows for faster response times and greater stability, as the data plane does not need to wait for real-time calculations from the control plane.
The control plane acts as the brain of the system, continuously optimizing routing weights based on a global objective function. This function minimizes end-to-end latency while accounting for capacity, health, and cache constraints. By having a global view, the control plane can anticipate bottlenecks and redistribute load before they become critical. This proactive approach is a significant improvement over the reactive nature of the feedback-loop system. The use of cached snapshots in the data plane implies a trade-off between routing precision and serving speed. While the weights may be slightly outdated, the benefit is a more stable and responsive system that can handle high volumes of traffic without degradation.
This architecture also enhances governance. The control plane’s decisions are based on explicit policies rather than emergent signals. This makes it easier to audit the system’s behavior and ensure that it aligns with organizational goals. The practical read is to map the policy language into controls, audit evidence, ownership, and reporting expectations for deployed AI systems [6]ai-redteam.comRouting LLM Inference in Production: From Engine… | AI Red TeamMy takeaway: Routing LLM Inference in Production: From Engine Signals to Policy — Qianru Lao & Lu Zhang, OpenAI is a governance signal. The practical read is to map the policy language into controls, audit evidence, ownership, and…Open source ↗. By having a clear separation of concerns, the system becomes more transparent and easier to manage. The control plane can be updated and optimized independently of the data plane, allowing for continuous improvement without disrupting service. The topic was featured in BestBlogs Daily on September 20, 2026, highlighting the broader industry interest in these architectural shifts [8]bestblogs.devBestBlogs Daily · 09-20 # LLM inference routing / MCP context engineering / AI security agents / open-weight adoption / diffusion inference [1] ★ Deep Dive · Routing LLM InferenceThe practical agent stack is routing, usable context, and security built in. The routing weights inside OpenAI's inference load balancer used to come out of a feedback loop.Open source ↗.
Compass Predictive Analytics

Global Optimization and Latency Minimization
The central goal of the new architecture is to minimize end-to-end latency. This is achieved through a global optimizer that considers all traffic across the fleet. The optimizer accounts for network delays, engine capacity, and cache states to determine the most efficient routing path for each request. This holistic approach ensures that the system as a whole performs optimally, rather than just individual components. The transition to a global optimizer suggests that OpenAI is prioritizing latency reduction as the primary objective, with safety and health as hard constraints.
The integration of penalties, capped retries, and load shedding further enhances the system’s stability. These mechanisms prevent the system from becoming overwhelmed during peak traffic periods. Load shedding allows the system to drop non-critical requests when capacity is limited, ensuring that critical requests are handled efficiently. Capped retries prevent the system from getting stuck in loops of failed requests, while penalties discourage routing to unhealthy engines. These features are essential for maintaining service quality in a dynamic environment.
The emphasis on latency minimization reflects the competitive nature of the AI industry. Users expect fast responses, and any delay can result in a poor user experience. By optimizing for latency, the system ensures that it can meet these expectations even under heavy load. This is particularly important for applications that require real-time interaction, such as chatbots or virtual assistants. The ability to provide low-latency responses is a key differentiator for AI providers, and the new architecture positions OpenAI to maintain its leadership in this area. The routing weights inside OpenAI's inference load balancer used to come out of a feedback loop, but now they are the result of complex, global calculations [5]youtube.comRouting LLM Inference in Production: From Engine Signals to Policy — Qianru Lao & Lu Zhang, OpenAI - YouTubeThe routing weights inside OpenAI's inference load balancer used to come out of a feedback loop. Engines reported signals, a controller smoothed them into a score, compared it to the fleet average, and nudged each weight up or down.Open source ↗.
Compass Predictive Analytics

Governance Implications and Safety Alignment
The shift from engine signals to policy has profound implications for AI governance. The move from emergent, signal-driven behavior to explicit, policy-driven governance requires a new approach to safety and compliance. Mapping this architecture to AI safety requires auditing the optimizer's objective functions and verifying alignment with human-intended utility and safety constraints. The control plane’s decisions must be transparent and explainable, allowing auditors to understand why a particular engine was chosen for a request.
The global view of the system also introduces new risks. A single point of failure in the control plane could disrupt the entire fleet. To mitigate this risk, the system must be designed with redundancy and fault tolerance. The control plane must be able to recover quickly from failures and continue to optimize routing without significant disruption. This requires robust monitoring and alerting systems to detect and respond to issues in real-time.
Furthermore, the explicit policies governing the system must be regularly reviewed and updated to reflect changing requirements. As the AI landscape evolves, so too must the governance frameworks that regulate it. The practical agent stack of routing, usable context, and security built in must be integrated into the governance model to ensure that safety is maintained throughout the lifecycle of the system. This includes ensuring that the routing decisions do not inadvertently compromise security or privacy.
The transition to a control-plane architecture also highlights the importance of cross-layer verification. As noted in recent research, static analysis of policy decisions embedded in each step is orthogonal to orchestration and currently left to ad-hoc Python functions. This gap must be addressed to ensure that the system’s decisions are safe and compliant. The integration of automated verification tools into the control plane can help to identify and correct policy violations before they result in harm [10]arxiv.orgFrom Inference Routing to Agent Orchestration Declarative Policy Compilation with Cross-Layer VerificationThese orchestration frameworks excel at their job—state management, fault tolerance, control flow—but static analysis of the policy decisions embedded in each step (which model to call, whether a tool invocation is safe, whether routing…Open source ↗. The practical agent stack is routing, usable context, and security built in, which must be supported by rigorous verification [4]x.comBestBlogs Daily · 09-20 # LLM inference routing / MCP context engineering / AI security agents / open-weight adoption / diffusion inference [1] ★ Deep Dive · Routing LLM InferenceBestBlogs Daily · 09-20 # LLM inference routing / MCP context engineering / AI security agents / open-weight adoption / diffusion inference [1] ★ Deep Dive · Routing LLM Inference in Production: From Engine Signals to Policy [Video]…Open source ↗.
Compass Predictive Analytics

Decisive Conclusion
The evolution of inference routing from feedback-loop load balancers to control-plane architectures represents a maturation of AI infrastructure. This shift is driven by the need for greater efficiency, stability, and governance. The new architecture allows for global optimization of latency while maintaining safety and health constraints. The separation of control and data planes enhances transparency and auditability, making it easier to align the system with organizational goals.
The implications of this shift extend beyond technical performance. They touch on the fundamental question of how AI systems are governed and controlled. The move from emergent behavior to explicit policy reflects a broader trend toward greater accountability and transparency in AI development. As the industry continues to scale, the importance of robust governance frameworks will only increase. The lessons learned from OpenAI’s IRB evolution provide a valuable roadmap for other providers seeking to optimize their infrastructure while maintaining safety and compliance.
The future of AI infrastructure will likely see further integration of policy-driven governance into routing and orchestration systems. As the complexity of AI workloads increases, the need for explicit, auditable policies will become even more critical. The transition from engine signals to policy is not just a technical upgrade but a necessary step toward responsible AI deployment. By embracing this shift, the industry can build systems that are not only efficient but also safe, transparent, and aligned with human values. The decisive move toward control-plane architectures marks a new era in AI infrastructure, one where governance and performance are seamlessly integrated.
Compass Predictive Analytics