Manual testing
How to test the parts of OneDrop that automated tests cannot reach: real agents, real sandboxes, real publishing.
How to test the parts of OneDrop that automated tests cannot reach: real agents, real sandboxes, real publishing.
Internal. These pages are hidden from the docs navigation and search. They’re for people working on OneDrop.
The Pest suite runs against a fake sandbox, a fake publisher and fake agents, so it can’t tell you whether Claude Code still signs in, whether a Shell survives a reload, or whether a published URL actually loads. These checklists cover that: the big concepts, tested by hand against real sandboxes and real AI providers.
| You changed | Run |
|---|---|
An agent version in docker/sandbox/Dockerfile, a runner or *Events class | Agents for that agent |
forwarder.mjs, HarnessRunner, the chat or the queue | Agents, all three |
start.sh, shell-keys.js, the Shell tab, onedrop-tool or docker/sandbox/lazy | Shell |
A Tools panel or the sandbox script behind it (db.php, auth.php…) | That panel in Tools |
A sandbox provider, SandboxUpdater, SandboxTools, snapshots | Sandboxes |
| A publisher, the gateway, hosting or sharing | Publishing |
| Before a release, or a big merge | Everything |
Each checklist names what must be true. When something fails, note the step and what you saw, then fix it with a Pest test that covers it where one can.
In .env, set SANDBOX_PROVIDER=docker, then build the image once and start everything:
php artisan sandbox:build-image
composer run dev
composer run dev runs the queue and the scheduler too. Without the queue the agent never starts; without the scheduler sandboxes never suspend or update.
dev@example.com / password. Use sam@example.com / password when a test needs a second person in the organization. Don’t reset the database to get them back without checking first; test against a copy instead.
In Settings → AI. Each agent needs its own connection:
| Agent | Needs |
|---|---|
| OpenCode | Any API key (OpenRouter is cheapest), or a ChatGPT sign-in |
| Claude Code | An Anthropic API key, or a Claude subscription sign-in |
| Codex | A ChatGPT sign-in, or an OpenAI API key |
DEV_CLAUDE_CREDENTIAL in .env seeds the dev user with a Claude connection.
“Make a page that says hello” is enough for most checks. Save bigger builds for the steps that need a real app, like Database or Hosting.
Every major feature gets a section here when it ships, alongside its spec, Pest tests and user docs (see AGENTS.md). Add it to the page for its area, or make a new page under docs/testing/ and add it to the Manual testing tab in docs/docs.json.
Write each test the same way:
### Publish to Tailscale (PUB-002)
1. Open a running project and click **Publish**.
2. Choose **Tailscale** and **Private**, then **Publish**.
- [ ] The status goes from **Publishing…** to a `*.ts.net` URL.
- [ ] The URL opens the app from a device on the tailnet.
- [ ] It doesn't open from a device off the tailnet.
noindex: true.