Skip to content
Methods and decisions

Data card

Data card: the Cradle seed and the live demo data

Last updated: October 2026 · Owner: Rin Huang

What it is

web/data/seed.db is a small SQLite database that every fresh Cradle database starts from: the local app, the tests, the evaluation sandbox and the production Turso database (created with turso db create project-cradle --from-file web/data/seed.db).

Composition

Table Rows What they are
users 19 The system inviter "Cradle", the two 2022 authors as founders, two shared demo accounts (mika, warden) and 14 fictional members
invitation_codes 21 One reusable Cradle code, 16 codes that someone joined with, 4 codes members are still holding
mails 5 The demo inbox for the two demo accounts: welcome mail, "someone joined with your code" notices and one sent invite
app_meta 2 Seed version and the date the seeded cohort closed

Join dates run from 7 January to 4 March 2022, the months the original project was built.

How it is made

  • Source: web/src/db/seed-data.ts (the people) and web/src/db/seed.ts (the rows), run by pnpm db:snapshot.
  • Deterministic: a mulberry32 generator with seed 20220107 draws every invite code and password salt, so the same source always produces the same bytes. A test pins the evaluation set's answers to this data.
  • Passwords: only the two demo accounts can sign in (credentials printed on /login), hashed with argon2id. Every other account stores a marker no hash scheme matches.

What is real and what isn't

  • Real: the two founders' usernames and names (Rin Huang, rNLKJA; Jiahong Zheng, kwitter777), who wrote the 2022 project; the 2022 rules and status messages.
  • Fictional: every other member, their names, occupations and made-up phone numbers. Email addresses use the reserved cradle.example domain.
  • Synthetic timings: every seeded code was created 26 hours before someone used it, and held codes are spread over the cohort's last week. Any statistic about timing on seed data describes this script, not people. /tree says so next to its estimates, and because every used code took exactly 26 hours it shows that median without an interval rather than a 95% interval of zero width.

Live data

On the production site, visitors can create accounts with an invite code and use the shared demo accounts. The site asks them not to enter real personal details, every address is in a demo inbox (nothing is emailed), and admins (including the public demo admin) can see all tables at /admin/records, with password hashes and session tokens hidden. Data on the live site is a demo artefact, not a sample of any population.

A reseed replaces members, codes and mail but keeps the append-only audit and AI logs. Member ids keep counting past every id ever issued, so a new member never inherits an earlier member's id or log.

Intended uses

  • Running and testing the app; a fixed dataset for the text-to-SQL evaluation.

Not suitable for

  • Any claim about how real communities grow, how people use invite codes, or how long they take. The cohort is too small (18 members) and fictional.

Known limitations

  • 18 members is too few for subgroup statistics, so intervals on /tree are wide where the data vary and absent where they don't (the synthetic timings). /tree presents them as a demo of the method on synthetic data, not as estimates about people.
  • Codes still open are right-censored: /tree counts redemption only for codes at least 7 days old and estimates time to redeem with a Kaplan–Meier median.
  • The structure (who invited whom) was written by hand to look plausible, so it has no realistic randomness.