Claude Code vs Codex vs Grok Bot (2026): An Advanced Comparison of Agents, Security and Cost
Claude Code (Anthropic), Codex (OpenAI, now inside the ChatGPT app) and Grok Bot (SpaceXAI, from the makers of Grok) are often mentioned together, but they are not the same kind of product. This guide is written for readers who already know how LLM agents work (tool use, context windows, sandboxes, MCP) and want the architectural and operational differences, not a feature checklist.
The short version
- Claude Code and Codex are coding-agent harnesses. They live next to your repository, edit files, run commands, open pull requests, and let you choose how much autonomy to grant. They are direct competitors.
- Grok Bot is a different category: persistent, always-on “AI teammates”, each with its own cloud computer, that sign in to your business apps and work through their interfaces (sales, ops, support, recruiting, engineering). It overlaps with the other two mainly in long-running, unattended work.
- The most important difference is where control lives: in a sandbox and permission system enforced by the harness (Claude Code, Codex) or in account-level authorisation plus instructions you give the Bot (Grok Bot, as far as its public pages describe).
At a glance
| Dimension | Claude Code | Codex | Grok Bot |
|---|---|---|---|
| Maker | Anthropic | OpenAI | SpaceXAI (Grok) |
| What it is | Agentic coding tool that reads a codebase, edits files, runs commands, integrates with dev tools | Coding agent for building features, refactors, migrations, code review; part of ChatGPT | Persistent AI teammates with their own cloud computer, for general business work |
| Where you use it | Terminal CLI, VS Code, JetBrains, Desktop app, web (claude.ai/code), iOS and Android, Slack, GitHub Actions and GitLab CI, Chrome | ChatGPT desktop app (Windows), IDE extension, CLI (npm), web/cloud, iOS | Desktop (macOS, Windows 10/11 x64) and iOS |
| Where the work runs | Your machine, Anthropic cloud VMs, or self-hosted environments (public beta) | Your machine (OS sandbox) or OpenAI-managed cloud containers | The Bot’s own cloud computer (runs 24/7, even with your laptop closed) |
| Project instructions | CLAUDE.md, auto memory; can also read AGENTS.md | AGENTS.md (global and nested, 32 KiB default cap), skills | Per-Bot memory and routines; account-level tools and skills |
| Extensibility | MCP, skills, hooks, subagents, Agent SDK | MCP, skills and plugins, subagents | Signs in to apps and sites like a person, including ones with no API or MCP |
| Scheduling | Cloud routines, desktop scheduled tasks, /loop | Automations: time-based; event-based on web (Gmail, Slack, GitHub PRs) | Routines on a schedule or triggered by events |
| Parallelism | Subagents, background agents, Projects (Desktop) | Worktrees, parallel cloud tasks, subagents | Many Bots in parallel; about 50 Bots per account, 6 per group chat |
| Entry price (vendor pages, USD) | Pro $17/month billed annually ($20 monthly); Max 5x $100; Max 20x $200 | Included in Free, Go ($8), Plus ($20), Pro (from $100); or API-key pay-as-you-go | Included with SuperGrok, Plus, Heavy and Cursor Pro, Pro+, Ultra and Teams plans; own usage allowance |
Prices are as shown on each vendor’s page on 23 September 2026, in US dollars, before tax, and can change. India pricing and availability were not verified.

Claude Code in depth
One engine, many surfaces. Anthropic’s docs say each surface (terminal, VS Code, JetBrains, Desktop, web) connects to the same Claude Code engine, so a repository’s CLAUDE.md files, settings and MCP servers behave the same everywhere. Sessions can be moved between surfaces: a cloud session can be pulled into a terminal with claude --teleport, and a terminal session can continue in the Desktop app.
- Context and memory: CLAUDE.md is read at the start of every session; “auto memory” saves learnings across sessions. If a repo already has AGENTS.md, Claude Code can read it alone or alongside CLAUDE.md.
- Extension points: MCP servers, skills (packaged workflows such as a PR-review command), hooks (shell commands before or after actions, for example formatting after every edit), subagents that a lead agent coordinates, background agents you watch from one screen, and an Agent SDK for building your own agents on the same tools.
- Autonomy in time: routines run in the cloud (they keep running with your computer off and can trigger on API calls or GitHub events); Desktop scheduled tasks run locally with access to your files;
/looprepeats a prompt inside a CLI session. - Permission modes: six named modes.
defaultreads only;acceptEditsadds file edits and common filesystem commands;planis for exploring before editing;autolets a separate classifier model review actions instead of you;dontAskdenies anything not pre-approved (for CI);bypassPermissionsskips checks and is meant for isolated containers and VMs only. On Pro, Max and Team plans the built-in starting mode is auto mode, and an organisation can disable it. - Isolation: a sandboxed bash tool with filesystem and network isolation; in Manual mode writes are limited to the launch folder and subfolders. Anthropic-hosted cloud sessions run in isolated VMs, limit network access by default, use a credential proxy so the sandbox never holds your real GitHub token, restrict git pushes to the working branch, and keep audit logs. Self-hosted environments (public beta) run sessions on your own infrastructure, where isolation and egress become your responsibility.
- Models: aliases such as
opus,sonnet,haiku,fable,bestandopusplan(Opus for planning, Sonnet for execution). Per the docs,defaultresolves to Opus 5.5 on Pro, Max, Team, Enterprise and the API. Default context is 200K tokens, with 1M available on many models; effort levels run from low to max. - Provider flexibility: the terminal CLI, VS Code and JetBrains also support third-party providers, so an enterprise can route through its own cloud account.
- Compliance: the docs point to Anthropic’s Trust Center for a SOC 2 Type 2 report and ISO 27001 certificate.
Codex in depth
The same agent, everywhere you code. OpenAI describes Codex as one agent across ChatGPT, the editor and the terminal, connected by your ChatGPT account: the ChatGPT desktop app (a Windows download is offered), an IDE extension, the CLI (npm i -g @openai/codex) and web/cloud environments, plus iOS.
- Instructions: AGENTS.md is discovered as a chain: global files in
~/.codex/first, then every level from the Git root down to the current directory, concatenated root-first so files closer to your working directory override earlier guidance. The default combined cap is 32 KiB (project_doc_max_bytes), so large repos need nested files or a higher limit. The chain is rebuilt on every run. Skills teach Codex a team’s standards and workflows. - Two independent controls: approvals and sandbox. Approval policies are
on-request(default in version-controlled folders),never(for CI) andgranular(keep some prompt categories interactive, reject others); the olderuntrustedmode has been retired. Sandbox modes areworkspace-write(default in version-controlled folders),read-onlyanddanger-full-access(discouraged). Network access is off by default..git,.agentsand.codexstay read-only even inside a writable workspace. - OS-level enforcement locally: Seatbelt on macOS, bwrap plus seccomp on Linux, and Windows Sandbox on Windows.
- Auto-review: setting
approvals_reviewer = "auto_review"routes eligible approval requests to an automated reviewer that looks for data exfiltration, credential probing, persistent security weakening and destructive actions. It can approve low and medium risk, denies critical risk and requires authorisation for high risk, and it adds extra model calls to your usage. - Cloud: tasks run in isolated OpenAI-managed containers. Setup can use the network to install dependencies; secrets are available only during setup and are removed before the agent phase, which then runs offline unless you enable internet access. Admins can control access with domain allowlists. Work can be started from the web, GitHub, GitLab, Linear or Slack and ends in a reviewable diff or pull request.
- Automations: on the desktop app, scheduled tasks run on your machine and need the app open; on the web they run in the cloud. Triggers can be time-based (including RFC 5545 recurrence rules) or, on web and mobile, event-based (Gmail, Slack, GitHub pull-request activity), but one task cannot combine event triggers with a schedule. Git worktrees isolate scheduled changes from your active work.
- Models and usage: the docs list GPT-6 Astra (most capable), GPT-6 Sol (complex coding and agentic work) and GPT-6 Luna (efficient, high-volume), with GPT-5.5 retiring on 14 October 2026. Usage is shared with ChatGPT Work and shown as estimated local messages per five-hour window; for example the pricing page estimates 15 to 150 GPT-6 Sol messages on Plus, and stresses these are not fixed limits and that cloud chats can use more.
Grok Bot in depth
Bots, not sessions. SpaceXAI’s design post describes a Bot as a persistent entity with a name, avatar, title, memory, runtime and tools, rather than a disposable chat. Each Bot has a dedicated cloud computer that can browse, manage files and run software, and you can watch it, pin a preview, or take over the screen when it needs help.
- Memory and routines belong to the Bot; tools and skills sit at account level. You can teach a workflow by asking a Bot to watch you do it once; it saves the routine and runs it itself next time. Routines can run on a schedule or on events, so work can start without you present.
- Multi-Bot coordination: Bots can message each other and share context in threads or group chats. Documented limits are about 50 Bots per account and six per group chat.
- Integration by interface, not by API: you sign a Bot in to your tools once and it uses apps and websites “just like you would”, including platforms with no clean API or MCP. That is a very different integration model from MCP tool calls.
- Availability: launched in beta in August 2026 for SuperGrok, SuperGrok Plus and Heavy, Cursor Pro, Pro+ and Ultra, and Cursor Teams Standard and Premium subscribers, with its own usage allowance separate from those plans. It is generally available for Enterprise, where SpaceXAI describes access, network and audit controls, a separate isolated environment per user, and no default permissions (Bots reach only accounts a user explicitly authorises). Compliance certifications were not stated in the announcement.
- Vendor case studies (self-reported): in procurement, a Bot reviewed about 125 vendors and flagged over $100,000 in savings, but needed explicit operator approval for any vendor-facing send and was barred from signing, purchasing or approving charges. In customer support, SpaceXAI reports a 175 percent rise in tickets handled and says 99 percent of refunds were processed without human intervention, after a phased “crawl, walk, run” rollout that began with internal notes needing human approval. These are the vendor’s own figures, not independent measurements.
- Model: the pages we read do not say which model powers Grok Bot; the SuperGrok listing mentions Grok 4.6, and SpaceXAI announced Grok 4.7 on 21 September 2026.
Head to head, for people who know the field
1. Where the guardrails are enforced
Claude Code and Codex both enforce limits outside the model: OS-level or container sandboxes, protected paths, network defaults, and approval policies. Both also add a model-based reviewer (Claude Code’s auto-mode classifier, Codex’s auto-review) on top, which is a second model judging the first, not a formal guarantee. Grok Bot’s public pages describe per-user isolated environments and “no default permissions”, and its case studies show limits written into the Bot’s instructions (for example, never sign or purchase). Instruction-level rules are useful but are softer than a sandbox, because they depend on the model following them. This is our reading of the published material, not something we tested.
2. Integration paradigm: API and MCP vs. the user interface
Claude Code and Codex reach the world through tools, MCP servers, CLIs and git. That is fast, auditable and scriptable, but only works where a tool exists. Grok Bot signs in to apps and uses them like a person, which reaches tools with no API but is slower to audit, depends on the app’s UI, and turns every login into a credential-scope question.
3. Autonomy over time
All three can work unattended, but with different failure surfaces. Claude Code routines and Codex web automations run in vendor clouds; Desktop and local scheduled tasks in both products run on your machine (Codex’s need the app open). Grok Bot’s Bots run continuously on their own computers. The longer the unattended window, the more you depend on approval gates, spending caps, logs and rollback (git for code; far less obvious for a Bot that sent an email).
4. Context and instruction hygiene
The two coding agents converge on markdown instruction files (CLAUDE.md and AGENTS.md, with Claude Code able to read AGENTS.md). Watch the size and layering rules: Codex’s 32 KiB default cap and root-first override order can silently drop late files in a big monorepo. Grok Bot keeps memory per Bot and skills per account, which is convenient but means each Bot’s “knowledge” lives in the product rather than in your repository.
5. Parallelism and orchestration
Claude Code offers subagents under a lead agent, background agents, Projects on Desktop and an Agent SDK. Codex offers worktrees, parallel cloud tasks and subagents. Grok Bot offers a Bot roster with group threads. For code, worktree or branch isolation matters; for business processes, shared state (the same CRM record edited by two Bots) is the harder problem.
6. Model choice and lock-in
Claude Code exposes model aliases, effort levels and third-party cloud providers, and supports a bring-your-own-account route. Codex ties to ChatGPT plans and models (GPT-6 family), with an API-key mode for CI that is pay-as-you-go and lacks the cloud features (such as GitHub review and Slack). Grok Bot is sold through Grok and Cursor plans and its model is not stated on the pages we read.
7. Cost shape
Claude Code and Codex are subscription-with-usage-windows for individuals and metered for API use. Codex publishes rough per-five-hour message ranges by model; Claude Code publishes plan tiers and says usage limits apply. Grok Bot is bundled with plans and has its own usage allowance. Because agent runs vary enormously in length, compare on cost per accepted change or completed task, not on the sticker price.
Risks that apply to all three
- Prompt injection: any agent that reads web pages, issues, emails or MCP output can be steered by hostile text. Prefer network-off defaults, allowlists, and human approval for outbound actions.
- Credential blast radius: Claude Code and Codex cloud both keep real secrets away from the agent phase (a credential proxy; secrets removed before the agent phase). A Bot logged in as you inherits whatever that account can do, so create narrowly scoped accounts for it.
- Silent partial failure: an agent can report success on work that is 90 percent done. Require tests, diffs or receipts, not just a summary.
- Supply chain: third-party MCP servers and plugins are code you are trusting; Anthropic says it does not security-audit MCP servers it lists.
- Vendor claims: savings, ticket and speed figures in launch posts and testimonials are self-reported.
How to choose
| If your main need is… | Look first at | Why (from the docs) |
|---|---|---|
| Terminal-first coding with fine-grained permission modes, hooks and an SDK | Claude Code | Six permission modes, hooks, subagents, Agent SDK, many surfaces, provider flexibility |
| Coding tied to a ChatGPT account, cloud tasks with offline agent phase, Gmail/Slack/PR-triggered automations | Codex | Approval and sandbox as separate dials, OS-level sandbox, cloud environments, event triggers on web |
| Non-coding, cross-app business workflows that run around the clock | Grok Bot | Own cloud computer, signs in to apps with no API, routines, multi-Bot threads |
| Strict compliance or on-premises needs | Check each vendor’s admin pages | Claude Code: self-hosted environments (beta) and Trust Center documents; Codex: domain allowlists and managed policy; Grok Bot: enterprise controls, certifications not stated in the announcement |
What we could not verify
- Benchmarks or hands-on quality: we ran none.
- Which model powers Grok Bot, and its compliance certifications.
- India-specific pricing, taxes and availability for all three.
- Real-world reliability of any auto-review or classifier feature.
Bottom line: treat Claude Code and Codex as two answers to the same coding question, and Grok Bot as a different answer to “who does the rest of the work”. The right pick depends on where you want control to sit and how much unattended autonomy you can safely audit.
General information only. Features, prices and limits change often. This is not an endorsement of any product, and it is not financial or security advice.
Sources (official, read on 23 September 2026)
- Claude Code documentation: overview, security, permission modes, model configuration, and product and pricing page
- OpenAI Codex product page; Codex docs: agent approvals and security, cloud, automations, AGENTS.md, models, pricing
- Grok Bot product page; SpaceXAI posts: Introducing Grok Bot, Designing Grok Bot, Grok Bot for Enterprise, procurement case study, customer support case study
