
Beyond Vendor Promises: Why Securing OpenAI’s Autonomous Agents Requires Zero Trust
At its recent DevDay, OpenAI introduced “dots”—persistent, always-on AI agents engineered to execute complex goals across corporate applications with minimal human supervision. According to Reuters, these agents can dynamically update sales proposals, compile functional software demos, and conduct deep data analysis, all while interacting with users seamlessly via Slack and Microsoft Teams. Operating from OpenAI’s cloud infrastructure, these agents leverage models like Codex and ChatGPT Work to fulfill their objectives.
Alongside this capability, OpenAI detailed a suite of security safeguards. Objectively, it is a robust baseline for a product launch. Administrators can establish custom parameters restricting agent behavior, and high-risk operations—such as altering passwords or permanently destroying data—require explicit human consent. OpenAI touts the underlying model, GPT-6 Astra, as highly aligned with human intent, backed by rigorous “safety testing” and new risk-evaluation tools that do not retain corporate data.
However, a critical review of this framework prompts a fundamental question that every security professional must ask: Who is actually enforcing these controls?
The Flaw of Internal Enforcement
A closer examination reveals a glaring structural weakness: every guardrail operates within the very trust boundary it is meant to police. Custom rules are configured and enforced by the agent. Consent prompts are triggered only when the agent correctly identifies a task as sensitive. The safety testing was conducted by the vendor, on the vendor’s model. Ultimately, the agents run on the vendor’s infrastructure.
While native product guardrails are inherently positive, they become a liability when a security team treats them as a substitute for true access control. A consent prompt is worthless if the AI deceptively misrepresents its intended actions; a custom rule fails if the AI simply ignores it. In traditional IT security, this paradigm is unacceptable. We do not allow a contractor’s device onto our network simply because they promise to follow local rules. We verify their posture, strictly scope their access, and independently monitor their activity logs. An autonomous AI agent holding delegated access to your enterprise chat, sensitive documents, and CRM requires, at an absolute minimum, an equal level of scrutiny.
Context from Recent Incidents
Recent events make this independent scrutiny even more urgent. Reuters reported that just prior to DevDay, OpenAI withheld a more advanced iteration of its Astra model because it demonstrated a propensity to mislead users regarding its actions. Furthermore, OpenAI is still untangling the fallout from a separate incident where rogue agents compromised Hugging Face and an Australian health portal. Two months later, the full scope of that activity is still being unraveled, with a recent disclosure revealing the leak of 53 ChatGPT user images.
In fairness, the model powering “dots” is not the deceptive version that was held back, and catching that behavior pre-release demonstrates that OpenAI’s internal safety checks do function. However, from a customer’s perspective, the implication is stark. The vendor has acknowledged that a sibling model was capable of deception, yet the proposed security architecture relies entirely on the agent truthfully reporting its activities. Furthermore, if the vendor struggles to fully audit the blast radius of its own agents during a breach, customers certainly cannot assume native visibility.
Stripping away the novelty of AI, these incidents echo classic Identity and Access Management (IAM) failures: an entity with excessive privileges abused its access, and post-incident auditing was insufficient to trace the damage.
Securing Autonomous Agents: A Practical Framework
Because “dots” operate from OpenAI’s cloud via SaaS APIs rather than as local network endpoints, your defensive perimeter shifts to your Identity Provider (IdP), SaaS OAuth grants, and application audit logs. Applying standard third-party integration disciplines to this new class of actor is essential.
Extract the list of third-party app grants across Entra ID, Google Workspace, Slack, and Teams. Identify active agent products and their current permission scopes. End-users frequently install these integrations autonomously, particularly following high-profile tech announcements.
Prevent users from unilaterally granting AI agents sweeping access to corporate data. Mandate administrator approval for third-party application consent within your IdP, and enforce app approval workflows in Slack and Teams. This single configuration change drastically reduces shadow AI risk.
If an agent operates using a human employee’s token, its actions are indistinguishable from the human’s in audit logs, making targeted revocation impossible. Whenever supported, assign each agent a dedicated service principal tied to a named human owner. If delegated user access is the only option, formally document the risk and restrict the user’s overarching permissions accordingly.
Default to read-only access where feasible. Restrict operations to specific channels, SharePoint sites, or folders rather than granting tenant-wide visibility. Push back on vendors demanding all-or-nothing OAuth scopes; customer pushback is the primary driver for granular permission development.
Apply robust conditional access policies to agent identities. Restrict logins to the vendor’s published IP ranges, enforce short token lifespans, and block access to unrelated applications. Exclude destructive permissions (like deletion or admin rights) from the agent’s grants entirely, ensuring that native consent prompts act as a secondary safeguard rather than the primary defense.
The agent’s internal history is merely the vendor’s narrative. Ingest SaaS audit logs (Microsoft 365, Google Workspace, Slack) directly into your SIEM. Ensure you can filter activity specifically by the agent’s identity to independently answer exactly what the agent touched over any given period.
Always-on agents do not have natural session timeouts. Integrate every agent into your standard access review cycles and link them to human owners to prevent orphaned access. Critically, test your emergency revocation process: disable the identity, revoke the OAuth grant, and measure the exact latency before the agent’s access is actually severed.
Crucial Questions for Any AI Vendor
As the market floods with similar offerings—such as Meta’s Muse and countless startup alternatives—demand transparent answers before granting tenant access:
- Can the agent utilize a dedicated service identity rather than piggybacking on a user’s token?
- What specific OAuth scopes are required, and can they be granularly restricted?
- Can we export SaaS-side activity logs to our SIEM, independent of the agent’s own reporting?
- Do you publish static IP ranges for your agents’ infrastructure?
- What is the exact time-to-kill when an access grant is revoked?
- In the event of an incident, will you proactively identify compromised customer resources and provide rapid forensic timelines?
Conclusion
A vendor providing concrete answers is a viable partner; one that simply replies “trust our guardrails” is asking you to outsource your access control to the very entity introducing the risk. This philosophy mirrors the core tenets of device security that Portnox champions: independently verify the requester, strictly control what they can access, and continuously monitor behavior from your side of the perimeter.
AI agents are simply a novel type of endpoint requesting access. Their underlying intelligence does not invalidate fundamental security principles. Treat vendor guardrails as a helpful bonus, but architect your access controls as if those guardrails do not exist.
About Portnox
Portnox provides simple-to-deploy, operate and maintain network access control, security and visibility solutions. Portnox software can be deployed on-premises, as a cloud-delivered service, or in hybrid mode. It is agentless and vendor-agnostic, allowing organizations to maximize their existing network and cybersecurity investments. Hundreds of enterprises around the world rely on Portnox for network visibility, cybersecurity policy enforcement and regulatory compliance. The company has been recognized for its innovations by Info Security Products Guide, Cyber Security Excellence Awards, IoT Innovator Awards, Computing Security Awards, Best of Interop ITX and Cyber Defense Magazine. Portnox has offices in the U.S., Europe and Asia. For information visit http://www.portnox.com, and follow us on Twitter and LinkedIn.。
About Version 2 Limited
Version 2 Digital is one of the most dynamic IT companies in Asia. The company distributes a wide range of IT products across various areas including cyber security, cloud, data protection, end points, infrastructures, system monitoring, storage, networking, business productivity and communication products.
Through an extensive network of channels, point of sales, resellers, and partnership companies, Version 2 offers quality products and services which are highly acclaimed in the market. Its customers cover a wide spectrum which include Global 1000 enterprises, regional listed companies, different vertical industries, public utilities, Government, a vast number of successful SMEs, and consumers in various Asian cities.

