# AGI House case — provenance

The interface on this case page is the real product, lifted from a private client work tree and
run against fixtures. This file records what was lifted, what was changed, and what deliberately
did not travel.

**Revised 2026-08-09.** The original version of this file listed the client's internal file paths
and the exact object counts of corpora that were never shipped. Those are the client's, not mine,
so they are gone. What remains is the method, which is the part worth publishing. A second
identity pass ran the same day and is described at the bottom; it exists because the first pass
was not good enough and the honest thing is to say so on the page rather than quietly re-upload.

## Screens

Each screen is the client's own HTML, CSS and JavaScript, with three classes of edit: data URLs
repointed at the local fixtures, remote dependencies (fonts, CDN emoji, remote logos, a GitHub
Pages data override) removed so the demo runs offline, and personal data replaced.

| Screen | Changes | Anonymisation |
|--------|---------|---------------|
| Case wrapper | Written for this portfolio | No personal data |
| People explorer | Local data, offline assets | Working set fully rewritten — see below |
| Event report | Local data; people iframe points at the local explorer | Anonymised applicants, events and projects; email lists replaced; sponsor brands kept, since a sponsorship is published marketing |
| Impact — people / companies | Remote v2 override removed; falls back to the v1 fixture, which is the only one that exists | Synthetic person and company labels, estimation prose scrubbed, dates to month precision, elevation numbers retained |
| Event shell | Local data | Same fixtures |
| Companies | Local data | Company names → sector labels; source URLs cleared |
| Events | Served from the catalog | Speakers, judges and organisers aliased; media URLs cleared |

## Anonymisation transform

- **Names** → deterministic aliases from a fixed vocabulary, so the same person is the same alias
  across every screen and the UI still demonstrates joins.
- **Emails** → `person+NNNN@example.test`.
- **Social and avatar URLs** → dropped.
- **Employers** → a sector-and-size label (`Foundation Models · 51-200`).
- **Score comments** → numeric summaries. The **numbers** are real; the free text that justified
  them named employers and GitHub identities and could not be published.
- **Dates** → month precision.

## Second identity pass (2026-08-09)

An audit of this case found the first pass had a gap wide enough to matter, and the fix is
recorded here rather than silently applied.

The gap: `headline` and its clone `ui.short_description` were never transformed. 111 of 312
headlines named a recognisable employer or school outright, and the rest carried a thesis topic
or a rare skill stack, which identifies a person just as well with one more search. Aliasing the
name above a fingerprint does not anonymise the fingerprint.

`scripts/scrub_identities.py` fixes it, and its design point is that headlines are **composed,
not filtered** — every published headline is rebuilt from fields that were already anonymised
(role score, sector label, catalog tags). A filter has a long tail by construction; a
reconstruction does not. The same pass:

- generalises the 47 org labels that came through the first pass with a real name still attached,
  including units small enough to identify one person on their own;
- restricts published tags to the curated catalog vocabulary, because the un-normalised free-text
  tags included bare company and university names;
- removes the venue street address;
- renumbers judge-sheet columns that had embedded judges' first names;
- aliases the people still hardcoded in two demo screens and stops their avatars resolving from
  real social handles.

It ends in a verification pass that re-reads what it wrote and exits non-zero if anything
identifying survives, so it is safe to re-run and cannot silently regress.

The anonymiser itself was also changed: it no longer contains the path to the client work tree,
nor the names of the individuals it removes. Both are supplied at run time (`AGI_SRC`,
`AGI_NAME_RULES`). A script that ships with a list of the people it anonymises has leaked them.

## Third pass (2026-08-09), and what the second one taught

The second pass checked a named list of fields. That is why it passed while five categories of
identifier were still shipping: it was looking where the previous problem had been. The third pass
replaces the field list with a sweep of **every string in every shipped fixture**, matched against
regexes for names, brands, street addresses, emails and handles, with a short allowlist of places
a brand is legitimately the subject — a technology stack, a public event title, a research
database, a city. Checking everything and listing the exceptions is a much harder thing to be
wrong about than listing what to check.

What that found, and what it says about anonymising by field:

- **Free-text captions carry what structured fields no longer do.** 483 timeline captions were
  written for humans and still read `Team joined <company>` long after the structured company
  field said `Company 041`. They are now composed from the event type, like the headlines.
- **Ordering is an identifier.** Numbering anonymised projects `01…34` in file order lets anyone
  holding the public demo list line the two up and recover all 34 names. They are now keyed on a
  hash of the name — stable across runs, unrelated to position.
- **An aggregate is only anonymous above a threshold.** The employer histogram had buckets of one.
  A single-person bucket labelled with an employer is that person. Labels are now sector
  groupings, and the chart says so in its own metadata.
- **Over-redaction is also a defect.** This pass initially stripped event titles and sponsor
  names, which are published marketing the client puts on a landing page. Removing them protected
  nobody and made the demo look like it had something to hide. They were restored deliberately.

Two fixtures also lost fingerprints that were never rendered but travelled in the file: publication
venues in the applicant records, and the four universities named in a cohort label.

Repairs to the demo itself, found while verifying the data still drew correctly — because a scrub
that leaves the page broken has not finished:

- `script.js` contained an `await` inside a non-async function, so the entire file failed to parse
  and **none** of the app's interactivity had been running. The enclosing function is now `async`.
- Fifteen references to three banner PNGs that live in the client's `generated/` directory and were
  never copied here. They now paint a CSS gradient, which is what the design's overlay expects
  underneath it, and removes both the failed requests and three project names from the network log.
- The optional v2 impact dataset was probed on every load and 404ed on every load. It is fetched
  only when an override names it.

## Explicitly not shipped

- The private bulk corpora. Their object counts are cited on the case page; the files were never
  copied here.
- Guest lists and dinner-guest analyses.
- A v2 impact file referenced by the client's deployment but absent from the source. It was not
  invented; the UI falls back to v1 and the page says so.
- Any credential, `.env` or MCP configuration. None were present in the lifted paths.

## Known imperfection, published rather than hidden

Anonymising the impact file minted synthetic company ids, which broke collisions the real ids
had. The private source resolves to 122 distinct companies; the fixture here resolves to 143. The
case page publishes 143 — the number you can recompute from the file you can actually see — and
names the discrepancy.

## Tooling

- `tools/anonymize_agi.py` — regenerates `app/` and `data/` from the read-only source tree.
- `scripts/scrub_identities.py` — the second identity pass, with its verification.

## Rebuild, 2026-08-30

The case page was rewritten from scratch. The iframe-of-fixtures presentation is gone from the
index; the page now shows screenshots of the production dashboards over the real database,
captured on 2026-08-30 with contact data removed before capture:

- Every screen was prepared with an injected script that deleted raw-JSON panes, contact and
  web-links sections, and masked any email-shaped text, then captured at 2× device scale.
- People shown by name are public founders and public figures (StackAI's co-founder, You.com's
  founder), shown with their own public record only.
- The workflow graph is rendered from the exported n8n definition (`wf-live.json`, 40 nodes);
  the prompt block quotes the live LinkedIn agent instructions verbatim; rubric weights are
  quoted from the production prompt text.
- Numbers were re-measured the day of the rebuild: 13,619 records with
  `step_8_2_detailed_score_prompt_completed: true` in the 449,758,138-byte enriched database;
  $1,861.47 summed from the provider transactions file.
- The broken interior screen `app/report-event-self-evolving-agent-build-day-20251101` was
  removed from the site.

The interactive app screens under `app/` are unchanged and still run on the anonymised fixtures
described above.
