Your Project Folder Gets a Librarian: Build Karpathy's LLM Wiki for Minutes, Specs and RFIs
A hands-on guide to Karpathy's LLM wiki pattern for architects: raw/ folder, schema, ingest, query and lint for minutes, specs and RFIs, versioned in git.
In 1945 Vannevar Bush imagined the Memex: a desk that would hold a person’s whole library, with trails linking one document to the next. On 4 April 2026, Andrej Karpathy published a GitHub gist, “LLM Wiki: A Pattern for Building Personal Knowledge Bases”, which picks up the thread Bush left open. According to the gist, the part Bush “couldn’t solve was who does the maintenance. The LLM handles that.”
The pattern is modest, and that is its strength. It is a filing discipline, not an AI product. Normally you search your raw documents again for every question. Karpathy proposes instead that a model should “incrementally build and maintain a persistent wiki—a structured, interlinked collection of markdown files that sits between you and the raw sources.” It has three layers: raw sources nobody edits, a wiki the model writes, and a schema file that holds the rules. It has three operations: ingest, query and lint.
Every architecture office knows where this lands. The project folder fills up with minutes, SIA phase notes, the spec and a stack of RFIs. Karpathy’s sharpest line belongs above every Projektablage: “The tedious part of maintaining a knowledge base is not the reading or the thinking—it’s the bookkeeping.” That bookkeeping usually falls to the project assistant who writes the Protokoll at 18:30, the quiet hero of every building that opened on time. This pattern does not replace their minutes. It makes them findable. The shared-drive search box has seen things.
←TODAY: In October 2026, a project’s memory lives in PDFs on a shared drive and in the heads of two people. →3012: Every building carries a plain-text history that any successor can open, cite and repair. Fulcrum: Markdown is boring enough to outlive the model that writes it.
The Tool
The Tool: The tool is Karpathy’s gist itself, deliberately abstract about directory structure, with every element described as optional. He names Obsidian as the reading interface and a git repository as the wiki’s home. Two community projects show how others have shipped it. karpathy-llm-wiki by astro-han (MIT licence) is an Agent Skill for coding agents with Ingest, Query and Lint commands; its author reports 94 wiki articles across 13 topics built from 99 ingested sources. llmwiki (Apache 2.0) turns the idea into a web app with an MCP server and a PDF/Office converter, plus a local mode on SQLite and the filesystem. Turning a gist into something you can install is real engineering, and both teams did it in the open. Today we build the bare version: folders, markdown, one schema and any coding agent that can edit local files.
Setup
Setup: One project folder, plus Git Bash or any POSIX shell.
mkdir -p 2611_Schulhaus_Musterstrasse && cd 2611_Schulhaus_Musterstrasse
mkdir -p raw/{minutes,specs,rfi,sia_phases,site_notes}
mkdir -p wiki/{decisions,people,components,open_issues,phases}
printf '# Index\n' > wiki/index.md
printf '# Log\n' > wiki/log.md
touch SCHEMA.md
git init && git add -A && git commit -m 'empty project wiki'
ls -R wikiThe rule of the house: raw/ is immutable, you drop files in and nobody edits them. wiki/ is the model’s territory, where you read and it writes. Convert PDFs to markdown with any converter, keep the original beside each one, and date-prefix raw files, for example 2026-09-18_minutes_bauherr.md.
First steps
First steps:
- Write the schema. Karpathy’s own example is a CLAUDE.md; name the file whatever your agent reads first. The operations and the two special files,
index.mdandlog.md, come from the gist. The page types and the SIA phase wording are our adaptation, not Karpathy’s — check your own SIA 112 phase list before relying on the numbers.You maintain wiki/ from raw/. Never edit raw/. Pages: decisions/ (what, who, date, source, SIA phase), people/, components/, open_issues/ (owner, due date), phases/. 1. Every claim cites [[raw/filename]] with date. No source, no claim. 2. If two sources disagree, write both, mark CONFLICT, link both. 3. Use [[wikilinks]]. No orphan pages. 4. Never invent dates, quantities or names. Write UNKNOWN. 5. German source text stays German in quotes. INGEST file: summarise, update affected pages, index.md, one log.md line. QUERY question: read index.md first, answer with citations, save if useful. LINT: report conflicts, stale issues, orphans, unsourced pages. - Seed it. Put 5–10 documents in
raw/: last month’s minutes, the current spec, the open RFIs. Tell the agent: “Read SCHEMA.md. INGEST every file in raw/, one at a time, and show me the log lines.” Karpathy describes a single ingest touching 10–15 wiki pages, so go one file at a time and watch where each document spreads. - Query. Ask a question that costs you ten minutes today: “Which decisions about the facade cladding are still open, and who owns them?” Then open two of the cited raw files and check them yourself.
- Lint and commit. Run LINT and read the report. Conflicts between the minutes and the spec are the payoff. Commit after every ingest.
The trade-off, plainly: the model can still misread a document, and the “no source, no claim” rule plus your spot-check are a safeguard, not a guarantee. noze.it’s write-up of 28 April 2026 describes the same three layers but reports no benchmarks, token costs or failure data, and nobody has shown how the pattern behaves on a 5,000-file project. Start with one project, one folder. One privacy line: client documents go to whichever model you run, so ingest only what your contract and the client’s NDA let you send to that provider (or use a local model), and strip personal data first.
From where I sit, the buildings that aged badly were the ones whose files went dark when a proprietary format died. Dated markdown, cross-linked and kept under git, is something a 25-year-old can still open decades from now.
Atelier
Atelier: For a 12-person studio with Archicad on every screen, the question is no longer whether AI can read the minutes but who owns the rules it files them by. Write the schema the way you would write a BEP chapter, and keep the decision about what may be ingested with the project lead, not whoever happens to run the agent. Monday move: on one live project, build the folder tree, paste in the schema, ingest only the last three sets of meeting minutes, and on Tuesday ask one real question with sources.
Hack
Hack: Hunt down every wiki page that makes claims without citing a raw file, then lock the clean state into git — your own lint pass, independent of the model. Any page the first line lists has broken rule one and goes back to the agent with “cite or delete”.
grep -rL '\[\[raw/' wiki --include='*.md' | grep -v -e index.md -e log.md
git add -A && git commit -m 'ingest 2026-09-18 minutes'
git diff HEAD~1 --statThe last line shows exactly which pages the latest ingest touched; git checkout HEAD~1 -- wiki/that_page.md rolls back a damaged one. A librarian that never takes a coffee break still gets audited.
Learn-it
Learn-it:
- The pattern: Andrej Karpathy — LLM Wiki: A Pattern for Building Personal Knowledge Bases
- Agent Skill implementation: karpathy-llm-wiki (astro-han)
- Web app with a local mode: llmwiki
- Walkthrough of the three layers: noze.it — LLM Wiki
- PAZ note: to tie the wiki’s phases/ pages to your Archicad project structure, PAZ Academy runs implementation sessions for offices adopting AI workflows.
Build the folder tonight, ingest three sets of minutes tomorrow, and make every answer show its source.
PAZ Kaffi · multidisciplinary editorial, led by PAZ Academy