CONTROL PLANE
The software and management layer that converts fragmented physical infrastructure into useful AI output. The research focus is not who has a dashboard; it is who can observe resources, make allocation decisions, recover failures, route workloads and ultimately monetise orchestration.
The five control layers
Current actors
StarFlow provides full-process training/inference management, heterogeneous-resource scheduling, topology-aware RDMA placement, observability and GPU fault self-healing. KC has therefore moved beyond raw compute rental into workload and compute management.
Smart Navigation is a facility-level “O&M brain” deployed across >90% of self-built sites, combining intelligent scheduling, cloud collaboration and autonomous operations. This creates a lower control-plane asset spanning physical capacity, cooling and energy.
Agentic Infra combines token production, context memory, runtime and unified general/AI scheduling. Its architecture increasingly joins physical and software planes.
Seed develops distributed training, high-performance inference and heterogeneous-hardware compilation. Volcano Engine exposes GPU and mGPU scheduling, topology and sharing controls.
A vertically integrated benchmark across models, cloud, workload scheduling, compute, storage and networking. It demonstrates what a full-stack control plane can look like when demand and infrastructure sit inside one platform.
MIIT's 1+M+N architecture explicitly calls for resource aggregation, selection, monitoring and interconnection scheduling across different subjects, architectures and regions. This is policy-backed market infrastructure, not yet proof of frictionless compute fungibility.
Capability matrix
| Actor | Workload | Compute | Network / storage | AIDC | Energy | Cross-provider? |
|---|---|---|---|---|---|---|
| KC | Strong | Strong | RDMA/storage integration | Limited | Limited | Not established |
| VNET | Limited | Emerging / investigate | Network + facility visibility | Strong | Strong facility-level | Not established |
| Huawei | Strong | Strong | Strong | Architecture layer | Digital Power / emerging integration | Within Huawei ecosystem strong; broader portability unresolved |
| ByteDance | Strong | Strong | Strong | Own campuses | Physical power footprint | Not established |
| Alibaba | Strong | Strong | Strong | Strong | Investigate | Not established beyond its cloud ecosystem |
| 1+M+N fabric | Market matching | Resource selection | Paths + monitoring | Indirect | Indirect | Explicit policy objective |
Where the value could migrate
Higher utilisation turns sunk accelerator capex into more billable output. Scheduling, sharing and failure recovery can therefore create value without adding silicon.
If software hides enough hardware differences, buyers can choose among more compute pools. That can reduce dependence on one accelerator source while increasing the value of the abstraction layer.
The largest white space is a trusted layer above individual providers that can discover, compare and route work across operators, regions and architectures. Policy is building some prerequisites; commercial ownership is unresolved.
VNET shows why the control plane extends below Kubernetes. AI-ready capacity depends on cooling, power, maintenance, reliability and site-level decisions.
CATL/DeepCtrls and VNET's energy-management work suggest a future boundary where compute scheduling and physical energy optimisation interact. Hyperscale dynamic coordination remains unproven.
If every provider remains a closed island, accelerator portability stays poor and national scheduling is mostly directory/trading infrastructure, the control-plane value pool will remain fragmented rather than becoming a dominant neutral layer.
Update · national coordination layer moves toward technical specification
The control-plane thesis should therefore be split into two interacting layers: provider control planes (Huawei, Volcano Engine, Alibaba, KC, telecoms and others) and an emerging national coordination layer concerned with discovery, identification, monitoring, scheduling, billing/trading and cross-region resource management. The unresolved value question is which functions remain public/common infrastructure and which become monetisable services for cloud, telecom and neutral infrastructure operators.
National standards register — adjustable-load project and related standards
Research priorities
Primary evidence
Kingsoft Cloud — StarFlow
Kingsoft Cloud — cloud-native AI suite
VNET — Innovation / Smart Navigation
Huawei — Agentic Infra
ByteDance Seed — Infrastructures
Volcano Engine — GPU scheduling
MIIT — national compute-interconnection nodes
MIIT — Compute Interconnection Action Plan