· 7 min read

By Argon Labs · Updated

A Disposable MongoDB Sandbox for Every AI Agent

Give each AI agent a MongoDB sandbox with Argon. Learn branch-per-agent isolation, dataset pins, TTL cleanup, and the capture requirements for undo.

AI AgentsMCPMongoDB

Current capability notes: checkout materializes a physical database; native actor attribution is per branch/run; undo needs complete images and retained history. CLI sandboxes require watch and scheduled sweep. Read the capability matrix.

The fastest way to lose a database is to hand an AI agent a write connection to production. The agent is capable and confident, and it will occasionally do exactly the wrong thing at full speed. The fix is not to keep agents read-only forever — it is to give each one a separate database for experiments. The agent can modify that branch while you review which changes to apply to its parent. Physical copies use storage; scope credentials to the intended database.

That is the branch-per-agent pattern, and Argon is built to make it a one-liner.

The branch-per-agent loop

  1. Fork a branch. Create a sandbox off production (or off a pinned dataset), optionally with a time-to-live. REST/MCP manage expiry while running; standalone CLI use requires a scheduled sweep.
  2. Let the agent work. Hand it the branch’s connection string. It reads and writes with any MongoDB driver — no code changes, no awareness that it is in a branch.
  3. Review the diff. When the agent is done, diff the branch against its parent to see exactly what it changed.
  4. Merge or discard. Merge the good work back as a reviewed data PR — or throw the whole branch away. Either way, production only ever sees changes you approved.

Because a fully captured, retained write range can be reverted, you also get a safety net after the fact: if something slips through, you can undo one branch actor’s captured changes. Separate agents need separate branches; undo can refuse a range when required history or images are incomplete.

Wiring it up

Start with the local setup. The MCP example also needs the project initialization on the agents page. Argon exposes the loop through these interfaces:

Over MCP (Claude Code, Cursor, any MCP client)

claude mcp add --transport stdio --env MONGODB_URI='mongodb://localhost:27017/?replicaSet=rs0' argon -- argon mcp

The agent gets 13 control tools over stdio, including sandbox creation, diff, merge, undo, snapshots, and pins. It uses a MongoDB driver with the returned connection for data reads and writes. See the MCP tool boundaries for the available interfaces.

From the CLI or CI

# Existing docs-review project; create a new agent-review sandbox
argon sandbox create -p docs-review --from main --name agent-review --ttl 1h
# Keep capture running before the agent writes
argon watch -p docs-review -b agent-review
# In a scheduled job, clean up expired CLI sandboxes
argon sandbox sweep -p docs-review

From Python (LangGraph, Mem0)

python3 -m pip install "argon-agents[langgraph]==0.2.0"

Install in a Python virtual environment; this uses the matching PyPI release. The package supplies a LangGraph checkpointer and a Mem0 sandbox factory. Prepare exact images on new collections before updates.

Review MongoDB documents before merging

A reviewed merge here means reviewing changes to database documents. It is separate from a GitHub or GitLab code pull request. Give two agents separate branches from the same retained pin, compare their proposed changes, and inspect the merge plan before applying it. A conflicting edit needs a resolution decision; creating a sandbox does not approve its writes. Follow the two-agent review example to inspect a conflict and the resulting document values.

Sandbox TTL versus MongoDB document TTL

MongoDB TTL indexes delete expired documents from a collection in the background. They do not replace a job that discards an entire agent sandbox, and expiration does not guarantee immediate deletion.

Argon sandbox TTL marks a branch for cleanup. REST and MCP manage that lifecycle while their service is running. With standalone CLI commands, schedule argon sandbox sweep for the project, as shown above, and inspect its result. A TTL value alone does not run the cleanup process. Finish any review or export you need before discarding the sandbox; keeping data for review and expiring it are different decisions.

Reproducible evals with pins

Evaluations are only meaningful if every run starts from the same data. A dataset pin is an immutable, named state of the database; each eval run forks a fresh branch from the pin, so the input is identical for that retained pin. Time travel is available within retained history. Pins protect the data they reference while the pin exists; keep backups and choose retention for your workload.

Why this matters now

Agents are moving from reading data to acting on it, and the database is where actions become permanent. Giving every agent its own branch turns an irreversible operation into a reviewable one — the same shift that pull requests brought to code. Production stops being the place agents experiment, and starts being the place their reviewed work lands.

Frequently asked questions

Why can’t an AI agent just use the production database?
Because the blast radius is your whole business. An agent that misreads an instruction can delete or corrupt live data, and you often cannot tell what it changed until later. A branch gives the agent a real database to work in while production stays untouched.
How does an AI agent get its own database?
With Argon, the agent (or your orchestration code) calls a single command or MCP tool to create a sandbox branch. It gets back an ordinary MongoDB connection string and reads and writes normally — no special SDK.
What is an MCP server for MongoDB?
Model Context Protocol (MCP) is a standard way to expose tools to AI agents. Argon ships 13 MCP tools for sandboxes, branch connections, diffs, merge plans, undo, snapshots, and pins. Document reads and writes use the returned MongoDB connection through a driver; retained-history queries use Argon’s CLI or REST API.
How do I make agent evaluations reproducible?
Use dataset pins: named references to captured document states. Fork each run from the same retained pin to control its starting document state. Matching input does not make agent outputs deterministic or guarantee identical physical database files.