Model card
Model card: Cradle's bring-your-own-key assistants
Last updated: October 2026 · Owner: Rin Huang · Related: DR-005, DR-008
Cradle does not train or host a model. This card describes how it uses third-party language models in two optional features, so the uses, limits and safeguards are written down in one place.
Models
| Default | Alternatives | |
|---|---|---|
| Provider | Anthropic | OpenAI |
| Model | Claude Haiku 4.5 (claude-haiku-4-5) |
Claude Sonnet 5.5 (claude-sonnet-5-5), or any OpenAI chat model id the visitor types |
| Who pays | The visitor, through their own API key | |
| Where the call runs | From the visitor's browser straight to the provider |
Claude Sonnet 5.5 calls ask for a low (notes) or medium (SQL) effort level and turn on Anthropic's default server-side refusal fallback; if the provider serves a request with a fallback model, the output's label and the AI log name both models, and in an evaluation run that answer is excluded from the configured model's accuracy.
Keys are stored per provider in the visitor's browser, and a key is checked against its provider before any request is built, so an Anthropic key is never sent to OpenAI or the other way round.
Intended use
- Invite note drafts (any member, /hub/invites). Suggest one or two friendly sentences to go with an invite code a member is emailing to a friend. The member inserts, edits or discards the draft; nothing is sent without them pressing "Send invite".
- Ask the records (admins, /admin/ask). Turn a question about the community ("How many members joined each month?") into one SQLite query, which the admin reviews and may edit before running it in a read-only sandbox.
- Evaluation runs (admins, /admin/ask/eval). Measure feature 2 against reference queries.
Out of scope
- Decisions about people: who may join, who is trusted, moderation. The models are never asked.
- Anything acted on without a person looking first. The one automatic step is the evaluation harness, which runs generated SQL against the frozen seed sandbox to score it; nobody acts on those results.
- Real personal data. Cradle is a demo with fictional members; visitors are asked not to enter real details.
Data sent to the provider
| Feature | Sent | Never sent |
|---|---|---|
| Invite note | The member's username, the chosen tone, an optional hint they type (up to 200 characters), and the instructions | The invite code, the friend's email, any other member data |
| Ask the records | The question, the sandbox schema (table and column names with one-line notes) and the instructions | Any row of data, query results, passwords, emails, phones, mail bodies, sessions, code strings |
| Evaluation | The reference questions, the schema and the instructions | The reference SQL and the expected answers |
What was sent and received is stored per call in the AI log (/ai-log), which members can export as CSV or JSON. The API key is never sent to Cradle's server. The log holds what the browser reports after each call; the browser checks first that the call can be logged, and an answer whose entry can't be written is not shown or used.
The demo admin account is open to every visitor and can read every member's log, which can't be deleted. On the shared demo accounts, free text a visitor types (an invite hint, a question, an edited note) and model prose that may repeat it are stored only as a fingerprint (length and a SHA-256 prefix); everyone else is asked not to type real names or details.
Training data provenance
The models were trained by their providers on data Cradle has no visibility into. See each provider's model documentation. Nothing from Cradle is used to train them through this integration beyond what the provider's own API terms state for API traffic.
Evaluation
- Design. 24 questions with hand-written reference queries (
cradle-sql-eval/v2; q03 was retired for having two fair answers and replaced by q25), run on a frozen copy of the seed data. A generated query scores when its result matches the reference: same rows, column names and order ignored, row order only when the question asks for it, numbers compared to four decimal places. - What counts. Rate limits, network or provider errors and unscored answers are excluded and counted separately rather than scored as wrong (transient ones are retried once), as are answers written by a fallback model. Refusals, truncated and malformed answers count as wrong. After a stopped run, configurations are compared only on the questions both reached. The quick run is a fixed 8-question subset covering every tag, and its results describe that subset.
- Uncertainty. Accuracy is reported with a Wilson 95% interval; median latency with a seeded bootstrap interval (2,000 resamples, seed 20220107); two configurations on the same questions are compared with the exact McNemar test and a paired bootstrap interval on the accuracy difference.
- Precision. With 24 questions, 17 correct (71%) has a 95% interval of 51% to 85%. Only large differences are detectable. Each question gets one generation, so the intervals describe which questions were asked, not how much a model's answers vary from run to run.
- Results. None published yet. The project has no budget for API calls, so results exist only when a visitor runs the harness with their own key; each run can be exported as JSON or CSV.
Known failure modes
These are the mistakes the questions were written to expose:
- Counting Cradle as a person. The system inviter is a row in
users; "how many members" should exclude it. - Time zones. Join times are Unix seconds in UTC. A member who joined at 9:40 am on 1 February in Melbourne joined on 31 January in UTC, so "joined in February" depends on the zone.
- Reusable codes. Cradle's codes are never marked used; questions about used or held codes must leave them out.
- Rounding and formatting. "Give a fraction" answered as a rounded percentage scores as wrong.
- Extra columns. Returning an id next to a username is reasonable to a person but fails execution match.
- Over-confident explanations. The model's explanation can describe the query it meant to write rather than the one it wrote. Read the SQL, not the summary.
- Notes that overpromise. A draft may describe the community in ways the member can't vouch for; the instructions forbid it, but drafts still need reading.
Safeguards
- Every output is labelled "AI-generated" with the model name.
- A person decides every time, and the decision (accepted, edited, rejected) is logged once and can't be changed afterwards.
- SQL runs only in an in-memory sandbox with analytic columns, read-only, one statement, no recursion (refused from the query plan too, since SQLite recurses without the keyword), no value-building or table-valued functions, a worst-case bound of 2,000,000 row visits read from the query plan, at most 200 rows back (DR-008).
- Rate limits per member on logging and on sandbox runs (per visitor as well on the shared demo accounts). A production Content-Security-Policy stops page scripts opening fetch, XHR, beacon or WebSocket connections to hosts other than this site and the two providers; it doesn't cover image loads or navigations, so on its own it doesn't stop a stolen key from leaving the page.
Ethical considerations
The main risks are misplaced trust (an admin reading a wrong number as fact) and data exposure. The first is addressed by showing the SQL, labelling the output and measuring accuracy with honest intervals; the second by never giving the model rows and never handling visitors' keys. The design is informed by the Australian Government's Policy for the responsible use of AI in government (DTA), the transparency principles in the EU AI Act and the NIST AI Risk Management Framework. It is not a compliance claim; Cradle is a personal demo, not a regulated system.