DR-008
Fail closed: fixes from the pre-release review
Status: Accepted · October 2026 · Author: Rin Huang · Supersedes the sandbox scan limit in DR-005 and the per-member demo limits in DR-007
Decision
Every way the review found for the AI features, the SQL sandbox or the estimates to fail open is closed in code, each with a test that reproduces the original problem, before the upgrade ships.
- Sandbox cost. A query runs only if a worst-case bound on the rows it can visit, read from SQLite's query plan and the sandbox's own row counts, is at most 2,000,000. Plans with a recursive step or a table-valued function are refused. This replaces DR-005's limit of three full table scans.
- The AI log fails closed. Before calling the provider, the browser asks the server whether one more call can be logged; if the entry can't be written afterwards, the answer is not shown or used. The log is described as what the browser reports, because that is what it is.
- One key per provider. Each provider's key is stored separately, and a key is checked against its provider (
sk-ant-for Anthropic, anything else for OpenAI) before any request is built. - Shared demo accounts. Free text typed there (an invite hint, a question, an edited note) and model prose that may repeat it are stored in the AI log as a fingerprint (length and a SHA-256 prefix). Rate limits on those accounts count per visitor IP as well as per account. A failed sign-in stores a keyed hash of a username that doesn't exist, never the text.
- Ids are never reused. A reseed no longer restarts the member id counter, and a member only sees AI-log rows written after their account was created.
- Evaluation. Provider, network and scoring failures, and answers written by a fallback model, are excluded from accuracy and counted separately. After a stopped run, A and B are compared only on the questions both reached. Question q03 is retired (gold set v2) and the quick run is a fixed subset that covers every tag.
- Estimates on /tree. Codes still open are treated as censored: redemption is "used within 7 days" over codes at least 7 days old, and time to redeem is a Kaplan–Meier median. An interval with no width is reported as such, not as a 95% interval.
Context
The career-aligned upgrade (DR-003 to DR-007) was finished and every CI gate passed. Before pushing, two independent reviews tried to break it on a production build. Between them they reported nine major findings, seven distinct problems once duplicates are merged, all reproduced:
- a 7-way self-join with index range filters passed the scan limit and kept the server busy for 26.5 seconds, during which an unrelated page took 25.7 seconds to load;
- an Anthropic key saved in the settings dialog was sent to api.openai.com after the visitor switched provider;
- the AI log failed open: when the log write was refused (for example by the per-account limit that every visitor of a demo account shares), the answer was still shown and used;
- text typed into the AI features on the public demo was readable by every visitor and could never be deleted;
- after a reseed, a new member could be given an earlier member's id and inherit their AI log;
- the evaluation harness scored rate-limited calls as wrong answers;
- /tree showed "26.0 h (95% CI 26.0–26.0)", an artefact of the synthetic seed timings, and counted open codes as never used.
Options considered
- Ship and list them as known limitations. Honest, but each one contradicts something /methods says the site does.
- Patch the reproduced inputs only (for example, refuse joins of more than three tables). Quick, but the next variation gets through.
- Fix each at its cause, with a test that reproduces the original failure (chosen).
- A hard statement timeout for the sandbox (run queries in a worker thread and stop it after two seconds). The strongest guarantee, but it means loading a native SQLite module inside a worker on Vercel, which I couldn't test locally; kept as the next step instead.
Why
- A bound that holds whatever the plan looks like. Each table access in the plan gets a factor f: the table's rows for a scan or a range search, the most rows that share one value for an equality search (times the longest literal
INlist), the work done so far for a scan of a subquery or CTE. Nested loops multiply and independent steps add, and either way the rows visited are at most the product of (1 + f) over the steps, minus one. Window functions can cost the square of their input, so the bound is squared when a query usesOVER. The limit keeps the worst case to a fraction of a second. - Fail closed, because the log is the governance claim. An AI answer the log doesn't contain is exactly what the log exists to prevent, so the answer goes rather than the claim.
- Separate keys make the mix-up impossible rather than unlikely. The prefix check is a second, independent guard.
- The demo is public on purpose, so what visitors type there must not become permanent, readable data.
- Infrastructure isn't the model. A 429 says nothing about whether a model writes correct SQL; counting it as wrong biases accuracy downwards and manufactures paired differences.
- Censoring is the textbook answer to "some codes haven't had their chance yet", and a zero-width interval from synthetic data is better reported as "no variation" than as certainty.
What happened
- The 7-way join now gets a bound of 20^7 = 1.28 billion rows and is refused at the plan stage in under a second. A 4-way version on the seed (bound 160,000; 130,321 rows) still runs, in about 1.5 ms. The 24 reference queries have bounds between 6 and 16,000.
- While fixing the sandbox I found a hole neither review reported: SQLite runs a self-referencing CTE recursively even without the
RECURSIVEkeyword, soWITH n(x) AS (SELECT 1 UNION ALL SELECT x + 1 FROM n) SELECT count(*) FROM npassed the keyword check and would never have finished. The plan shows aRECURSIVE STEP, which is now refused; a test pins it. - The key mix-up was reproduced with a mocked network (
Authorization: Bearer sk-ant-…sent to api.openai.com). The test now asserts that no request is made at all. - With 3 of A's 8 calls rate-limited, the harness used to report A at 5/8 (62.5%, Wilson 30.6–86.3%) and "B vs A +37.5 pp", although A was right on every question it answered. It now reports A at 5/5 with 3 excluded and no difference.
- /tree now shows "26.0 h" and says that every used code took exactly that long. The seeded share is unchanged at 16/20 (80.0%, Wilson 58.4–91.9%) because every seeded code is older than 7 days; on the live site, newly made codes no longer pull it down.
- A test reproduces the reseed problem (member id 20 handed out twice).
data/seed.dbis byte-for-byte unchanged, because the seed already used explicit ids. - Q03 ("How many invite codes have been used to join so far?") had a fair second reading: the founders joined with Cradle's reusable code, giving 17 rather than 16. It is retired, never reused, and replaced by q25, which names member-issued codes.
- Costs: one extra round trip before each AI call (the quota check); demo visitors behind one shared IP still share a budget; the log is still reported by the browser, so the server can't prove that every call was reported. The bound is conservative, so a legitimate query on a much larger live table could be refused.
What I'd change
- Add the hard timeout (option 4) on top of the bound.
- Have the server issue a log id before each call, so a call that starts but is never reported shows up as missing.
- Run each evaluation question several times, to measure how much a model's answers vary from run to run; today the intervals describe question sampling only.
- Show the per-question bound next to "Run" in Ask the records, so an admin can see why a query was refused before sending it.