Engineering·6 min read
Inside the build → test → fix → deploy loop
Generating code is the easy part. Here is how MoboStudio AI proves that what it built actually works.
Most AI builders are organised around one arrow: prompt → code. MoboStudio AI is organised around a loop — build, test, fix, deploy — and the architecture exists to make that loop safe, observable, recoverable and billable.
1. Claim, reserve, plan
A run starts as a queued job. A worker claims it with a conditional update, so a redelivered job is simply skipped. Before any model is called, credits are reserved. Then a reasoning model produces a plan through a single tool call: a summary, an ordered task list and — for new projects — a name. The plan appears in the interface as a live checklist.
2. Code with real tools
A coding model works through a tool loop with a fixed budget of turns. Its tools are deliberately small and explicit:
// Tools available to the coding agent
list_files · read_file · write_file
edit_file // requires a unique, exact match
delete_file · run_command · update_task · finish`run_command` is intentionally narrow: one command per call, no shell. The binary must be npm, npx, node, pnpm or yarn; pipes, redirects, chained commands and global installs are rejected. Commands get a timeout and a sanitised environment that never includes platform secrets.
3. Verify like a user would
When a project runtime is available, the result is installed, built with npm run build, started as a dev server and opened in headless Chromium. The check collects console errors, page errors, failed network requests and a screenshot. Anything found goes back to the coder — for a limited number of rounds.
If it still can’t solve the problem, the agent stops, explains what it found and asks. It never burns credits in an endless loop.
4. Version, settle, ship
The run records the version before and after, so a build’s exact changes can be reviewed — and any file undone — later. Credits are settled at actual cost, never above the reservation. Deploys must pass a health check before they go live; a failed deploy never replaces a working one.
Why it’s built this way
- The API never executes generated code; expensive work always goes through queues.
- Every step has a stop condition: turns, attempts, credits, context.
- Every change is attributable to a run, and reversible.
None of this is visible when things go well. It’s what makes the result trustworthy when they don’t.