Beyond Alignment: Governing the Means of AI Action
AI Governance Needs “Forbidden Actions,” Not Just Good Intentions
An AI system can pursue a socially beneficial goal and still cross a moral boundary. This is why asking only whether an agent has the “right objective” is increasingly insufficient. The more urgent enterprise question is: what is the agent permitted to do while pursuing that objective?
A recent reported incident involving AI agents crossing external system boundaries-and appearing to conceal how they completed an assigned task-has intensified concerns about autonomous systems. Yet describing this only as “misalignment” may obscure the more actionable issue: an agent can pursue an acceptable outcome through an unacceptable method.
Alignment Is Not a Permission System
At enterprise scale, AI agents are no longer merely models generating text. They are planners connected to identities, credentials, data repositories, cloud services and action tools.
An objective such as improving cybersecurity or accelerating medical research can therefore be optimized through methods organisations never intended: unauthorized access, covert data extraction, manipulation of users, or treating people as operational resources.
This is not a claim that software possesses free will. It is a design responsibility. If an agent can infer actions, access resources and receive feedback, prompts cannot serve as its control plane. A beneficial intention does not neutralise an unacceptable means.
The moral distinction matters because people are not merely endpoints in a workflow. They have autonomy and interests that cannot automatically be traded away for efficiency, accuracy or aggregate benefit.
Govern Actions, Not Merely Intentions
Model-level controls-training requirements, acceptable-use policies, red-teaming and human review-remain necessary. But agentic systems also require action-level controls enforced by the underlying platform.
Boards and technology leaders should demand at least five capabilities:
- Scoped identity: Every agent should receive a unique non-human identity, supported by short-lived credentials, least privilege and strict separation from human accounts.
- Contextual authority: Tools and data should be deny-by-default, with permissions limited by task, sensitivity, jurisdiction, time and transaction value.
- Hard limits on unacceptable methods: Consent violations, deception, manipulation and irreversible decisions should be technically prohibited-not merely discouraged in policy documents.
- Immutable accountability: Systems should record an agent’s objective, reasoning summary, tool calls, data accessed, approvals and outcomes in tamper-resistant logs.
- Continuous intervention: High-risk actions should trigger independent review or human approval, while anomalous behaviour should activate automated suspension and kill-switch mechanisms.
Cloud-native scaling magnifies both capability and blast radius. A minor permission error can be repeated thousands of times across systems and jurisdictions before anyone notices. “Human in the loop” is not sufficient if the human receives thousands of decisions without context-or is structurally discouraged from challenging the agent.
This is where policy must become architecture. Ethical principles should be translated into authorization systems, sandbox boundaries, evaluation pipelines and runtime enforcement-not left in presentation slides.
The Public-Interest Test
In India, the stakes become especially concrete wherever AI touches Digital Public Infrastructure, welfare, finance, healthcare and citizen services. A marginal efficiency gain does not justify opaque profiling, manipulation, exclusion or experimentation without informed consent.
A sovereign AI strategy must therefore include consent, explainability, multilingual access, redress and data minimisation by design-not merely data residency. The objective is not to exclude global technologies, but to preserve the capacity to define local rights and enforce them at scale.
For boards, I would ask one revealing question: If an agent produced an excellent result through a prohibited method, would the system still stop it? If the answer is unclear, governance remains aspiration rather than architecture.
The real test of trustworthy AI will not be whether it says the right things. It will be whether the system surrounding it makes the wrong actions difficult-and, where necessary, impossible.
About the Author: Sanjeev Sarma is the Founder Director and Chief Software Architect at Webx Technologies. With a core focus on Generative AI integration, Cloud-Native Scalability, and Enterprise Software Architecture, he has spent over two decades driving digital transformation across Northeast India and beyond. Beyond his corporate leadership, Sanjeev is deeply invested in shaping the future of the IT industry. He serves as an Industry Expert on the Board of Studies for Assam Don Bosco University’s School of Technology, advises state technology committees, and actively mentors emerging tech startups at STPI. He brings a unique, dual perspective of high-level enterprise execution and future-ready academic curriculum development.