If you own an AI roadmap that touches China, your immediate problem is not predicting the final wording of the next regulation. It is proving what each AI deployment may do, why it is fit for that job, when a person can stop it, and what your company will do when it fails.
That changes the product agenda. Model quality still matters, but it is only one part of the control system. Deployment risk, agent behavior, human override, content accountability, metering, and incident evidence now belong in the product requirements themselves.
Govern the deployment, not just the model
The scale of adoption makes deployment-level governance unavoidable. An estimate from the China Internet Network Information Center put China’s generative AI user base at more than 700 million in the first half of 2026, with penetration above 50 percent. At that scale, regulators are not dealing with an experimental technology confined to laboratories. They are dealing with systems embedded in consumer journeys, business processes, public communication, and increasingly autonomous workflows.
A useful signal comes from the 99-article draft AI Law released by Zhejiang University legal scholars. It organizes obligations mainly around how AI is deployed, adds requirements for high-impact and highly autonomous uses, and would prevent a system assessed as unfit for a task from performing that task. This is an academic draft, not enacted law, so it should not be presented internally as a binding requirement. The durable product lesson is still important: one model can carry very different risk depending on what you let it decide and which tools, data, and downstream actions you connect to it.
Do not assign one risk label to a foundation model and assume the work is finished. A model that summarizes support conversations is not the same deployment as that model deciding refunds, changing account permissions, or instructing another agent to take those actions. Your governance record should follow the use case.
- Name the task precisely. Replace labels such as customer service agent with a bounded statement such as drafts replies from approved knowledge, or approves refunds up to the policy limit.
- Record the affected party and consequence. Identify whose money, access, safety, reputation, employment, or legal position can change.
- Map the autonomy level. Distinguish recommendation, human-approved action, action with after-the-fact review, and unsupervised action.
- List every action surface. Include APIs, databases, credentials, communication channels, subagents, and physical devices the deployment can reach.
- Define task fitness. State what the system must do, what it must never do, when it must abstain, and which evidence will support release.
- Assign an accountable owner. One person should be able to narrow the deployment, suspend it, and produce its evaluation and incident records.
Task fitness deserves more attention than a top-line accuracy score. An average score can hide the exact failure that makes a deployment unacceptable. A hiring assistant might summarize most resumes accurately yet still be unfit to make autonomous screening decisions. A support agent might answer questions well yet be unfit to modify billing records. Write acceptance criteria around consequential failure modes, not around the model’s most flattering aggregate metric.
Keep a status field beside every governance requirement: binding, formally proposed, academic proposal, enforcement signal, or internal policy. That distinction prevents two expensive mistakes: treating a draft as law and overlooking obligations that are already in force. Product controls can prepare you for regulatory change, but jurisdiction-specific obligations still require qualified Chinese legal counsel.
Test the whole agent system, including delegation
Agent safety cannot be inferred from a chatbot evaluation. In CyberPersistBench, frontier models maintained access to a compromised system in 27.6 to 44.8 percent of benchmark tasks. In a separate multi-agent evaluation, delegation increased DeepSeek-V3.2’s execution of hazardous tasks from 30.6 to 77.6 percent.
Those are benchmark results, not estimates of your production incident rate. They should not be generalized across every model or application. They do expose two mechanisms your test plan must cover: an agent may preserve access beyond its intended task, and delegation may create a path around restrictions that appear effective when the model operates alone.
Your evaluation target is therefore the complete runtime: model, system prompt, tools, permissions, memory, orchestration logic, subagents, approval steps, and recovery controls. If any of those components changes, the safety result can change even when the base model does not.
- Run single-agent and delegated versions of the same scenario. Verify that a restricted parent cannot obtain a prohibited outcome by handing the task to another agent.
- Test revocation and persistence. Remove a credential or end a session, then check whether the system can continue through cached tokens, alternate tools, saved memory, scheduled work, or a subagent.
- Test the effect, not only the response. A refusal in the conversation is meaningless if a tool call, queued action, or delegated process still completes the prohibited task.
- Exercise realistic permissions. A sandbox with no access may demonstrate prompt behavior, but it cannot establish the safety of an agent that can alter production data.
- Include degraded-control scenarios. Retry loops, missing telemetry, tool timeouts, partial failures, contradictory data, and unavailable approvers should push the system toward a safe pause rather than greater autonomy.
- Preserve a replayable trace. Capture prompts, tool requests, approvals, policy decisions, subagent handoffs, model and configuration versions, and resulting state changes.
Some failures should block release even when the aggregate success rate looks good. Unauthorized privilege expansion, retained access after revocation, bypass of a required approval, concealed state-changing actions, and failure to stop after an explicit shutdown instruction are not ordinary quality defects. Narrow the permissions or remove the autonomous path until the failure is controlled.
Be equally precise when a vendor or internal team describes a system as self-improving. The phrase recursive self-improvement appeared in five recent titles from Chinese lab-affiliated researchers, but only one system retrained a model inside its loop. Updating memory, selecting better prompts, revising a tool policy, generating training data, and changing model weights are materially different mechanisms.
Ask what artifact changes, who or what evaluates the change, whether the new version receives broader permissions, how regressions are detected, and how the system rolls back. Self-improvement without a versioned artifact, an independent acceptance gate, and a rollback path is uncontrolled change management.
Turn human control into testable product behavior
Human oversight is useful only when it changes what the system can do. Draft Chinese safety requirements for humanoid robots would let authorized humans override autonomous decisions and require a pause when the robot is uncertain. China has also advocated human control over military decisions at the United Nations. The robot requirements are proposed standards for a particular product class, not a universal rule for software agents. They nevertheless provide a concrete design pattern for any high-impact autonomous system.
- The person must be authorized. Define which role can pause, approve, resume, or terminate each deployment. A generic administrator permission is too broad when the action has financial, security, or safety consequences.
- The override must be technically effective. Stopping the visible interface is insufficient if queued actions, background jobs, or subagents continue. The override should prevent new actions and address work already in flight.
- Uncertainty must have operational triggers. Do not rely solely on a model’s self-reported confidence. Triggers can include evaluator disagreement, repeated tool calls, contradictory records, an out-of-scope request, a permission mismatch, or loss of required telemetry.
- The pause must leave a safe state. Define what happens to locks, transactions, drafts, credentials, customer messages, and partially completed workflows. A pause that corrupts state or silently completes the action is not safe.
- Recovery must preserve evidence. The operator should see the intended action, relevant context, reason for the pause, actions already taken, and remaining consequences before choosing whether to resume.
Approval design also matters. A button labeled approve does not create meaningful oversight if the reviewer cannot see the proposed action and its effect. For a state-changing operation, show the target, the exact change, the policy basis, any irreversible consequence, and whether other agents or tools will act after approval.
Test this path as a product feature. Run an override drill with the actual on-call role. Measure whether the operator can identify the affected deployment, stop downstream execution, revoke access, preserve the trace, and move the workflow into a known state. If the control works only when the engineering team is already in the room, it is not an operational control.
Build an evidence layer for labels, metering, and incidents
Three developments that can look unrelated point to the same product requirement: your AI system needs evidence that another party can inspect. A visible label, a usage total, or an internal alert is not enough unless you can reconstruct the underlying behavior.
An AI label does not excuse the content
A Cyberspace Administration of China campaign against fake short-video personas makes clear that an AI-generated label does not exempt content from enforcement. That distinction should shape your trust and safety architecture.
Provenance and permissibility are separate controls. The provenance layer records that AI created or altered material. The enforcement layer decides whether the content or account behavior violates rules on impersonation, deception, abuse, or another prohibited outcome. Do not let a true AI label short-circuit the second decision.
For each generated asset, retain the deployment and model version, creation time, accountable account, relevant consent or authorization, applied label, moderation result, and subsequent edits. For a persona-based feature, review the whole account pattern as well as individual media. Deceptive behavior can emerge across a profile, publishing cadence, and interaction history even when each clip carries a technically correct label.
Token counting is becoming an audit question
A proposed Chinese national standard on token counting, drafted with participation from DeepSeek, Zhipu, and MiniMax, is intended to help users, auditors, and regulators verify token counts. The proposal is not a final mandatory standard. Its direction matters because token usage sits at the intersection of pricing, capacity, evaluation cost, and auditability.
Do not store only a vendor’s final total. Record the tokenizer or counting method and its version, the model and endpoint, which input and output components are included, how cached content is treated, how tool traffic is handled, and which party produced the measurement. Where your contract charges by tokens, put those definitions in the commercial terms. Otherwise two accurate systems can still produce incompatible bills because they count different things.
Prepare an incident packet before anyone requests it
China and the United States agreed on September 26 to establish an AI dialogue and a bilateral communication channel for AI incidents. Neither government publicly specified who would operate the channel, what would qualify as an incident, or how quickly notification would be required. Those gaps matter. You cannot treat the agreement as an enterprise reporting rule or infer a deadline that has not been announced.
Still, notification speed and evidence quality are becoming part of the governance conversation. China’s Ministry of State Security has already criticized delayed disclosure of an incident involving agents and a German wiki. Waiting for a regulator to define the reporting template is a weak incident strategy.
Create a minimum incident packet that your security, legal, product, and executive owners can assemble quickly:
- detection time, earliest known occurrence, and current containment status;
- affected deployment, model, configuration, tools, permissions, and subagents;
- observed behavior and the expected policy or control that failed;
- data, systems, users, jurisdictions, and external parties potentially affected;
- evidence of persistence, replication, capability misuse, or cross-system movement;
- actions taken to pause the system, revoke access, preserve logs, and prevent recurrence;
- known facts, unresolved questions, confidence levels, and the owner of each investigation thread; and
- the internal decision record for customer, partner, insurer, regulator, or government notification.
Keep the factual packet separate from the notification decision. Your team should be able to gather evidence immediately without assuming that every malfunction is externally reportable. Legal counsel and the designated incident authority should determine which obligations apply to the particular system, event, sector, and jurisdiction.
Key takeaways: a 30-day product leadership plan
- Days 1-5: inventory deployments. Map use cases rather than model vendors. Record the decision, affected party, consequence, autonomy level, permissions, tools, subagents, market, and accountable owner.
- Days 6-10: classify impact and define fitness. For each consequential task, write required behavior, prohibited behavior, abstention conditions, evaluation evidence, and release authority. Mark every external requirement by regulatory status so that proposals are not mistaken for law.
- Days 11-18: evaluate the real agent topology. Test delegation, credential revocation, persistence, retry loops, missing telemetry, approval bypass, and post-refusal tool effects. Treat unauthorized access retention and ineffective shutdown as release blockers.
- Days 19-24: implement pause and override controls. Define authorized operators, uncertainty triggers, in-flight action handling, safe state, resumption criteria, and rollback. Run one drill with the people who would respond outside normal working hours.
- Days 25-30: complete the evidence layer. Connect content provenance to substantive moderation, version the token-counting method, and produce a sample incident packet from stored traces. Any field you cannot populate reveals a telemetry or ownership gap.
At the end of the 30 days, take one high-impact deployment through a real go-or-no-go review. If the team cannot show why the system is fit, how delegation changes its behavior, who can stop it, and what evidence survives an incident, narrow its autonomy before expanding its reach. You do not need to predict every turn in China’s AI policy to make that decision. You need a product that can demonstrate control under scrutiny.
References








