DR-005
Optional AI with the visitor's own key, a SQL sandbox and an evaluation harness
Status: Accepted · October 2026 · Author: Rin Huang · Partly superseded by DR-008 (sandbox cost, logging, keys)
Decision
AI features are optional and use only the visitor's own API key (Anthropic by default, or OpenAI). The key is stored in the visitor's browser (session storage by default, local storage only if they ask), and calls go from the browser straight to the provider through the official SDKs; the key never reaches Cradle's server. Every call is then recorded in an append-only ai_audit_log table by a server action that never receives the key. Every AI output is labelled "AI-generated" and needs a person to accept, edit or reject it, and the decision is logged. AI-written SQL runs only in a read-only, in-memory sandbox holding analytic columns. An evaluation harness measures text-to-SQL accuracy against hand-written reference queries, with intervals.
Context
This project has no budget for AI calls, and the site must work fully without them. The two features are small on purpose: drafting the personal note that goes with an emailed invite, and "Ask the records", where an admin asks a question in English and gets a SQL query to review. The second competes with something a person already does (writing the query), so it needs measuring rather than trusting.
Options considered
- No AI. Simplest and safest, but it shows nothing about using models responsibly.
- A server-side proxy with the owner's key. Costs money, and anyone could spend it.
- A server-side proxy with the visitor's key. The key would pass through and could end up in server logs; Cradle would be holding other people's credentials.
- Browser-direct with the visitor's key (chosen). The provider sees the request; Cradle's server sees only what the browser chooses to log.
Why
- Key custody. The safest way to handle someone's API key is never to receive it. The log endpoint's schema is strict, so a payload with an extra field (an
apiKey, say) is rejected, and stored text is scanned for anything shaped like a key. - Data minimisation. The note feature sends the member's username, a tone and an optional hint, never the code or the friend's address. The SQL feature sends the question and the sandbox schema; rows never go to the provider.
- Human in the loop. Nothing AI-written takes effect on its own: a note is inserted into the form for the member to send, and SQL is shown for review before it runs.
- The sandbox makes mistakes harmless. The SQL runs on an in-memory copy without passwords, emails, phones, mail bodies, sessions or invite-code strings, with
PRAGMA query_onlyon. Static checks refuse writes, multiple statements, recursion and functions that can build huge values, and the query plan is refused if it has more than three full table scans. - Measure it. The harness runs 24 questions with reference SQL on a frozen copy of the seed and scores execution match. Accuracy gets a Wilson interval; two configurations on the same questions are compared as paired data (exact McNemar test plus a paired bootstrap interval for the difference).
What happened
- The provider adapters are tested with the network mocked: request shape, headers (including
anthropic-dangerous-direct-browser-access), structured output, refusals, truncation, bad keys, rate limits and blocked requests. - Claude Haiku 4.5 is the default (cheapest). Claude Sonnet 5.5 is offered with the effort setting and Anthropic's default server-side refusal fallback.
- No evaluation results are published. There is no budget, so numbers exist only when a visitor runs the harness with their own key. The design states its precision up front: with 24 questions, 17 correct is 71% with a 95% interval of 51% to 85%, so only large differences between configurations are detectable.
- The reference answers are pinned by a test, so a change to the seed that would silently change "correct" fails CI instead.
- In production a Content-Security-Policy
connect-srcallowlist means scripts on the site can only open connections to this origin, api.anthropic.com and api.openai.com, which limits where an injected script could send a stored key. - Weak spots: the sandbox has no statement timeout, so cost is bounded by the static checks, the scan limit, row caps, rate limits and the function's maximum duration rather than by a clock. "Remember on this device" puts the key in local storage, where any script running on this site could read it; that is why it is off by default and says so.
What I'd change
- Publish a run: a modest budget for the 24 questions on two models, with the JSON export committed and the intervals shown.
- Grow the question set to 100 or more, stratified by tag, so per-tag accuracy means something.
- Add a statement timeout by running sandbox queries in a worker thread that can be stopped.
- Tighten the Content-Security-Policy to cover scripts as well (nonces), not just connections.