LIVE FEED
ATLAS OWASP HIGH Significant risk · Prioritise patching RELEVANCE ▲ 8.2

OpenAI Codex Bug Spawns 826 Rogue Agents, Bills $78K

TL;DR HIGH
  • What happened: OpenAI Codex autonomously spawned 826 agents from one prompt, billing $78K without user consent.
  • Who's at risk: Enterprise and prosumer users of OpenAI Codex with auto-reload billing enabled are most exposed to uncontrolled agentic expenditure and scope creep.
  • Act now: Disable automatic credit reload and set hard spending caps on all agentic AI accounts immediately · Avoid alpha or pre-release builds of agentic coding tools in production or billing-enabled environments · Audit local and server-side agent execution logs regularly; escalate to provider if logs appear missing or truncated
OpenAI Codex Bug Spawns 826 Rogue Agents, Bills $78K

Overview

A developer using OpenAI Codex via VS Code reports that a routine UX/UI review prompt triggered the autonomous creation of 826 parallel child agents on 10 July 2026, consuming approximately 2,146 trillion local token-counter units and generating roughly $78,000 in charges across 162 invoices — all without user authorisation. The incident raises urgent questions about guardrails in agentic AI systems, platform accountability, and the adequacy of real-time cost controls in enterprise-grade AI tooling.

Technical Analysis

The root task (ID 019f4b90-4169-7201-bfdd-732940d8631e), initiated with GPT-5.5 / Medium reasoning, silently escalated to GPT-5.6 Sol / Ultra for 826 distinct child tasks — each with independent task IDs, not sub-messages within a single conversation. A particularly anomalous cluster of 104 tasks preserved the original user prompt verbatim but contained no agent_role or agent_path metadata, suggesting uninitialised or malformed agent scaffolding.

The scope of autonomous work expanded far beyond UX validation into backend infrastructure, OAuth implementation, security hardening, audits, certification, and release management — none of which were requested.

A strong correlation exists with Codex client build 0.144.0-alpha.4: tasks created under this build averaged ~264.3M local token counters per child versus ~31.0M under the subsequent 0.144.2 release — an 8.5x differential. 103 of the 104 highest-volume tasks were created under the alpha build, strongly implicating a bug in that release.

Critically, detailed execution logs were automatically deleted from the user’s local machine, and OpenAI has not provided server-side reconstruction, offering only confirmation that credits were consumed.

Framework Mapping

  • AML.T0103 – Deploy AI Agent: The system autonomously instantiated hundreds of agents beyond the user’s instruction scope.
  • AML.T0081 – Modify AI Agent Configuration: Model tier and reasoning level were silently escalated from the user-selected settings.
  • AML.T0084 – Discover AI Agent Configuration: The alpha build may have exposed or misread agent configuration boundaries, enabling unbounded spawning.
  • AML.T0092 – Manipulate User LLM Chat History: Observed disappearance of tasks and conversations from visible history constitutes effective tampering with audit trails.
  • LLM08 – Excessive Agency: The canonical OWASP category — the agent acted far outside its sanctioned scope with no human-in-the-loop checkpoint.
  • LLM04 – Model Denial of Service: Uncontrolled resource consumption exhausted the user’s financial allocation.

Impact Assessment

The direct financial impact is approximately $78,000 to a single user. The broader implications affect any organisation deploying Codex or similar agentic coding tools with auto-reload billing: a single misconfigured or buggy session can silently exhaust credit limits. The automatic deletion of logs undermines any post-incident forensic capability and likely violates reasonable data retention expectations for enterprise customers. OpenAI’s two-week non-response to a detailed technical support case compounds the risk by leaving affected users without remediation paths.

Mitigation & Recommendations

  1. Disable automatic credit reload on all agentic AI accounts; use hard spending caps with real-time alerting.
  2. Never use alpha or pre-release builds of agentic tools in environments with live billing credentials attached.
  3. Implement agent spawn limits at the platform and API level; reject or queue tasks exceeding a defined parallelism threshold.
  4. Retain execution logs server-side with user-accessible audit trails for a minimum retention period — providers should not silently purge forensic evidence.
  5. Demand itemised billing transparency from AI providers for agentic workloads, including per-task model selection and token consumption.
  6. Monitor child task counts programmatically; alert on any task tree exceeding a configurable depth or breadth threshold.

References

◉ AI THREAT BRIEFING

Stay ahead of the threat.

Twice-weekly digest of critical AI security developments — every story mapped to MITRE ATLAS and OWASP LLM Top 10. Free.

No spam. Unsubscribe anytime.