December book

Wega Labs · Review edition · 2026-08-02

December

The complete working book for a little artificial world that keeps living: the dream, the causal machinery, the research contract, every open decision, and the audits that tried to break it.

35source documents
58,490source words
≈ 4h 15mcomplete read
Gate 0still open

Before page one

The reader’s desk

This is a faithful compilation, not a replacement specification. Every chapter names its source file and retains the source text. This desk exists to make the decisions and inconsistencies visible before more of the world is built.

The most important status collision. The roadmap says implementation must remain unstarted until Gate 0 closes. The README says Phase 1 is underway, and the deterministic event/RNG/replay spine already exists. We should formally call that foundation a bounded pre-gate spike, or revise Gate 0 and the roadmap to match what has actually happened.

The seven decisions that need you

  1. Budget: maximum cash spend after existing credits, emergency cutoff behavior, and the canonical-world versus research-ensemble allocation.
  2. Licensing: permissive-only reuse or deliberate compatibility with reciprocal licenses.
  3. Deployment: canonical machine and OS, and whether cloud infrastructure is in scope.
  4. Pace: the initial relationship between real and simulated time.
  5. Publication: private/local history, public streaming, or a staged combination.
  6. Research posture: accept or revise the claims ladder, R0 protocol, and separation of canonical exhibit from registered cohorts.
  7. Naming: reserve December for this living world and distinctly rename or separate December Sato.

Randomness and human messiness

The design allows chance, uneven personalities, destructive motives, glory-seeking, revenge, altruism, martyrdom, irrationality, domination, risk, and mistakes. It forbids a system-wide reward for viewer engagement or a hidden storyteller creating war because war would be exciting. Individuals may want terrible or theatrical things; the world itself cannot secretly want a better episode.

Reading paths

One-hour alignment: Overview, pass-2 audit, risks/open questions, roadmap, and ADRs. Full understanding: Parts I–IV in order, then the audits. Implementation: focus on chapters 06, 09, 10, 11, 14, 15, 16, and the ADRs.

No chapter contains that phrase.
Part I

Enter the dream

The promise, the public story, and the present state.

OrientationOVERVIEW.md
2,837 words
Inside this chapter

December#

A world that keeps living.

An 11-minute introduction. The full specification is in wiki/; this is the version you read first. A shorter version is on the web at wegalabs.com/plan.


In one paragraph. December is a small settlement of artificial residents that keeps running whether or not anyone is watching. Each resident is played by a language model that chooses what to attempt; a simulation decides what is actually possible, so nobody invents food, occupies two places at once, or knows something they were never told. Every event is recorded with its cause, so any outcome can be traced back through the decisions, observations, and chance that produced it. Nothing is running yet — the design is finished and audited, the foundation is built, and no resident exists.

Here is why it is built that way.


Eighteen people live in a valley. They farm, hunt, argue, build, make promises, break them, fall ill, have children, and die. Nobody is watching most of the time.

You come back after two days and something has happened. A dry season thinned the stream; the council rationed water; two households held grain back; a child's testimony changed a vote; the losing steward left with four followers and camped at the weir upstream. Three people were hurt in the confrontation. The people who nursed them caught what they had.

The point is not that this story is dramatic. The point is that you can ask why, and get an answer made of evidence rather than narration. Every clause in that paragraph resolves to a recorded event with a cause, a timestamp, an actor, and — where chance was involved — the exact random draw that produced it. You can walk backwards from the injury to the confrontation to the vote to the rainfall. You can ask what any resident knew at any moment, and be told only what they could actually have known.

That property, and not the story, is what December is for.

1. The problem with agent societies#

Put several language models in a room, give them names and backstories, and they will produce something that reads like a society. They will negotiate, form alliances, betray each other, and write you a very good chronicle of it afterwards.

The trouble is that none of it is load-bearing. Language models are extraordinary at sounding like people and hopeless at being ledgers. Ask a model to narrate a world and grain appears from nowhere, a character is in two places on the same afternoon, a wall gets built in an evening, and a war starts because a war was the interesting thing to write next. Nothing is conserved, so nothing is at stake.

The deeper failure is epistemic. When such a society produces a war, you cannot tell whether it happened because the harvest failed or because "war" was simply a plausible next paragraph. The history is unfalsifiable. It cannot surprise you, because it could have gone any way at all, and you would have believed that too.

Everything below follows from taking that problem seriously.

2. One commitment: the simulator owns truth#

December has a typed, deterministic kernel that owns all authoritative state — every gram of grain, every person's location, every claim on a field, every hour of the day. The kernel is not clever. It does arithmetic, checks preconditions, and refuses things.

Language models never touch it. A resident's model receives a private observation packet — what that person can see, remember, and has been told — and proposes an action in a fixed schema. The kernel validates the action against reality and either performs it, rejects it, or returns what would be possible instead. Prose describes the world. It never determines it.

This sounds like a limitation on the models. It is the opposite. A choice only means something if it could have failed. When a resident decides to give away food during a shortage, that decision is interesting precisely because the food is finite, the ledger is real, and the cost will arrive later whether or not anyone narrates it. Remove the constraint and the generosity is just a nice sentence.

The same commitment produces the causal graph. Because every change to the world is an event with recorded parents — this rainfall caused that soil moisture, this observation caused that decision, this rule authorised that transfer, this random draw produced that infection — you can reconstruct the reasoning for any outcome without asking a model to speculate about it afterwards. When the system says someone died of a wound infection following a fight over water, that is a query result, not an interpretation.

There is a read-only narrator, which we call the director. It writes summaries, ranks what mattered, and points a camera. It cannot change anything, and any claim in its summaries that isn't backed by a cited event is rejected before you see it. If it ever needed the power to make things happen in order to keep the world interesting, the project would have failed on its own terms.

3. What eighteen people can and cannot be#

The settlement begins with twelve adults and six dependants. This number is not a gameplay choice; it is small enough that every person can be simulated with real interiority and large enough for households, factions, and disagreement.

It is also, honestly, below every published threshold for a viable isolated human population. Simulations of forager demography put the floor around forty people and near-certainty around a hundred and fifty. With twelve breeding adults, inbreeding accumulates at roughly four percent per generation — four times the rate conservation biologists consider tolerable — and by the third generation every child is more closely related than the offspring of first cousins.

And yet. Pitcairn Island was founded in 1790 by twenty-seven people and still exists. Tristan da Cunha was founded by about fifteen and still exists, carrying a thirty-six percent asthma rate two centuries later as the price. Rapa Nui fell to roughly a hundred and ten people in 1877 and recovered.

So the honest position is that eighteen people is marginal rather than impossible — and the interesting question is what makes the difference. The historical record answers it consistently: not headcount, but connection. The colonies that failed — Norse Greenland, the Polynesian settlements on Henderson and Pitcairn — died when their networks were severed, not when their numbers were low. Greenland's western settlement, six to eight hundred people, could be emptied in two centuries by about one departing household a year. No catastrophe required.

This reshapes the design. The neighbouring group is not primarily a source of raids and disease; it is the settlement's lifeline, and it has to exist early rather than late. Partner exchange, in-migration, and a seasonal regional gathering are load-bearing mechanics. A world where the neighbours only ever show up to fight would be a hazard generator wearing a costume.

Scale honesty cuts the other way too, and this is where most simulations quietly cheat. Consider disease. The number of people required to sustain an acute immunising infection like measles — the critical community size — is somewhere between two hundred fifty thousand and five hundred thousand. December has eighteen. Four orders of magnitude short.

So a conventional epidemic model, run at this scale, does one of two things: it fades out immediately and contributes nothing, or it gets tuned until epidemics appear at a satisfying rate. The second option is fabricated epidemiology wearing the costume of rigour, and it would be very easy to ship.

Instead the health model is rebuilt around what can actually persist in eighteen people: chronic and latent infections, where the host is the reservoir; environmentally transmitted disease, where the reservoir is the water; zoonoses, where the reservoir is an animal; and parasites, which accumulate with sedentism. Acute epidemics still happen, but only as rare introductions from outside contact, and they behave the way virgin-soil epidemics actually behave — either the infection fails to take hold at all, or it takes nearly everyone. There is no average epidemic at this scale, and any implementation producing one has a bug.

That correction made the design better, not smaller. The persistent disease pressure now comes from the settlement's own decisions about water, waste, and storage. Two populations living in the same ecosystem can differ tenfold in parasite load depending on whether they stay put. The settlement's success at storing grain is what brings the rats.

4. Why the seed grain matters more than the government#

If you want to know where December expects its drama to come from, it isn't the constitution. It's a ratio.

In medieval European agriculture — the best-documented premodern cereal record, tens of thousands of manor-year observations — a field returned about three and a half to four grains for every grain sown. That means roughly twenty-seven percent of every harvest has to be taken away from hungry people and put back in the ground. Contemporary estate managers used a three-to-one return as their break-even line, which tells you how close to the edge the whole arrangement ran. In something like one field-year in fourteen, the yield is bad enough that half the harvest must be withheld.

Sit with what that does to a settlement in a bad autumn. Eating the seed corn is available, immediately rational, and catastrophic on a twelve-month delay. Nobody has to be villainous. The trap is arithmetic.

Hunting works the same way and produces the opposite social result. Measured return rates among Hadza hunters have a coefficient of variation above eight: a hunter targeting large game comes home with nothing on more than ninety-seven percent of days. The settlement eats meat because someone succeeds, never because anyone reliably does. That is not a colourful detail — it is the reason food sharing, reciprocity, and obligation networks exist at all. A simulation that hands hunters their average return has deleted the reason for society while appearing to model hunting correctly.

Grinding grain by hand runs around seven-tenths of a kilogram an hour, which for a household means three to five hours every single day, forever, falling almost entirely on women in every ethnographic case on record. If you model fields and harvests but not processing, you have silently handed the settlement several free person-hours a day and misrepresented who is doing the work.

The pattern in all three: the institutions we hope to see emerge should be forced by material facts, not offered as a menu. There is no democracy module. There is a bounded grammar for making rules, offices, and obligations, and there are pressures that make rule-making worth the trouble.

The clearest evidence that this works comes from Palmerston Island, settled in 1863 by four people — one man and three women, two of whom were cousins. It produced over a thousand descendants. It did that by organising itself into three exogamous branches with marriage inside a branch forbidden: an invented kinship rule, transmitted and enforced across generations, that solved a genetic problem the founders could perceive but not name.

That is precisely the shape December is built to produce. A material constraint — a shrinking pool of unrelated partners — generating a durable institution that nobody designed in advance. If a settlement independently arrives at an exogamy rule under demographic pressure, that will be a more convincing result than any election, because the pressure is measurable, the mechanism is legible, and the settlements that fail to invent it are visible in the same experiment.

5. What we refuse to claim#

A project like this attracts overclaiming, so the limits are worth stating as plainly as the ambitions.

This is not a model of real human societies. It is a fictional valley with invented parameters, many borrowed from populations that have nothing to do with it. Nothing December produces is evidence about how people actually behave, and no result should ever be used to argue about real communities.

Simulated conflict rates carry no evidentiary weight. The scholarly literature on violence in small-scale societies spans an order of magnitude, and much of that disagreement turns out to be definitional rather than empirical — the two most-cited papers on opposite sides do not even report the same quantity, and six of the same societies appear in both with opposite conclusions. December's conflict rate is a design choice inside a contested range. It cannot show that violence is natural, and it cannot show that peace is.

We are not studying consciousness, and we are not studying whether a person could survive being copied. Those questions motivate the work and are not tested by it. A resident who maintains a coherent identity across a model swap has demonstrated something about persistent synthetic agency. It has demonstrated nothing whatsoever about experience. The distance between those two claims is the whole distance, and we do not intend to blur it.

A behavioural finding might be a fact about the model provider. If residents rarely escalate a conflict, that could be the world working — or it could be the safety training of whichever model is playing them. So refusal rates are measured and reported alongside every behavioural result. A conflict statistic without its refusal denominator is uninterpretable.

Reproducibility has a precise meaning. Given the same code, configuration, and recorded model responses, a history replays to identical state. That works because every quantity in the world is an integer — grams, millilitres, seconds — so nothing depends on how a particular machine rounds a decimal. No language model provider offers reproducible sampling, so the recorded responses are part of the permanent record rather than a cache; a replay that cannot find one fails loudly instead of quietly inventing a different history.

6. How you would know if it worked#

The hardest problem in a project like this is not building it. It is telling the difference between a world that produced something and a world that was arranged to produce it.

Three habits guard against that.

The founding state is declared, randomised, and varied. An earlier draft of the scenario described the founders as disagreeing about property, leadership, and risk, with land claims left conveniently ambiguous. That is not a starting condition; it is the plot, pre-installed. If the settlement later fractures over property and elects someone to sort it out, we would have observed only what we planted. So the initial world specifies quantities, locations, relationships, and capabilities — never dispositions toward a future conflict — and it is generated from a recorded seed so that experiments can vary it. Randomising does not make it unauthored: the distributions, the constraints, and the seeds we accept are all research decisions, and they are written down.

Every claimed cascade is measured against a null model. With eighteen people, births and deaths are single digits per decade, so variance rivals the mean and a meaningful fraction of worlds will simply die for no interesting reason. Unless a collapse is compared against the same demography with no social mechanisms at all, "the faction dispute destroyed the settlement" is indistinguishable from "a settlement of eighteen people collapsed, and there happened to be a faction dispute."

Labels are defined before the runs, not after. Words like faction, war, and governance form need structural definitions — readable off the institution registry, with persistence requirements — fixed in version control before the experiments that use them. A definition chosen after seeing the results is not a finding.

Beyond that, the criteria are ordinary: histories should differ across seeds; removing a mechanism should measurably change outcomes; residents with different values and histories should make measurably different decisions rather than becoming one personality wearing twelve names; and a week of unattended running should require no human repair.

And there is one clean failure condition. If the world needs a storyteller to stay interesting, it has failed. Quiet seasons are allowed. A world that is only compelling when someone is secretly arranging events is a very expensive way to write fiction.

Where it stands#

The design is complete and has been through two independent adversarial audits, the second of which corrected the first on several points, including some it had gotten confidently wrong.

Implementation has started at the foundation rather than the scenery. What exists today is the part that cannot be retrofitted: integer-valued world state, a tamper-evident chain of events, reproducible random streams, an append-only history, and a replay path that reconstructs the world exactly. It runs on a deliberately trivial world — grain moving between containers — because the point of a foundation is to be tested before anything is built on it.

There is no ecology yet, no people, no cognition. Nothing is alive.

The nearest milestone is a headless world with weather, water, crops, and demography, running thousands of times with no language models involved at all, to find out which settlements survive and why before any resident is asked to make a single decision.


Details, sources, and the reasoning behind every choice above are in wiki/. The audits are in AUDIT-FINDINGS.md and AUDIT-FINDINGS-PASS-2.md, including the parts where they disagree.

Current statusREADME.md
1,141 words
Inside this chapter

December: The Living Terrarium#

Status: design revised after two audit passes; Phase 1 underway
Site: wegalabs.com · the argument, in 6 minutes
Implementation status: the determinism spine is built and tested (src/, tests/). No ecology, no residents, no cognition — nothing is alive yet.
Audit status: Gate 0 remains open. See the first audit and independent pass 2. The second pass accepted the engineering premise. The unsafe public human-participant flow has since been removed; versioning, owner decisions, preregistration, parameter provenance, and scope decisions remain open.

December is a plan for an always-running, observable settlement whose inhabitants can survive, form relationships, build, govern, trade, disagree, elect leaders, split into factions, fight, migrate, reproduce, and die. The ambition is not a scripted AI soap opera. It is a small causal world capable of producing histories that surprise us for reasons we can inspect.

As a Wega Labs research program, its credible near-term subject is persistent synthetic agency: what allows an artificial resident to remain identifiable and causally continuous across memory consolidation, model changes, bodily and social change, and counterfactual forks. Consciousness and human continuity beyond biology are long-horizon motivations, not questions the current terrarium can decide. See the lab charter and claims ladder.

The recommended first world is a fictional early agrarian valley. Its limited technology keeps the model tractable while still allowing scarcity, seasons, construction, property, commons, kinship, ritual, law, elections, disease, migration, raiding, war, and extinction. The current design cohort is 12 cognitively rich adults, 6 lightweight dependants, and a neighboring band represented at lower resolution until contact makes individuals relevant; Gate 1 demographic ensembles determine whether that number survives as the canonical choice.

The core promise#

When the observer returns after two days, something consequential may have happened—but not because a narrator rolled “war” on a story table. A drought may have reduced stream flow, damaged crops, depleted stores, intensified a property dispute, empowered a faction, caused a failed raid, spread infection among the wounded, and left the settlement near extinction. Every link is recorded. Recorded-event replay reproduces state exactly; kernel re-execution is deterministic only within its declared pinned environment and manifest.

This is an observable AI terrarium, not a claim that we have recreated humanity. “Realistic” means causally coherent, materially constrained, behaviorally plausible within declared limits, statistically tested, and auditable. It does not mean scientifically validated human prediction.

Reproducibility has three meanings: portable reconstruction from recorded events, deterministic re-execution of the kernel within its declared numeric environment, and fresh counterfactual simulation. Only the first is an event-log replay; a branch with new model responses is a new experiment. See determinism and replay.

Start here#

DECEMBER-BOOK.html — the entire product, research, architecture, ADR, and audit corpus in one searchable, self-contained review edition. It includes per-chapter notes, read tracking, and Markdown feedback export.

OVERVIEW.md — an 11-minute introduction in prose. What December is, the handful of ideas that actually matter, and what it refuses to claim. Read this before the wiki. A shorter version is on the web at wegalabs.com/plan.

The code is in src/december/. Phase 1 begins with the parts that cannot be retrofitted — integer-valued world state, hash-chained events, reproducible RNG streams, and exact replay — running on a deliberately trivial world so the foundation is tested before anything is built on it. pytest runs the determinism suite from wiki/14 §D9.

The documents below are a specification, not an introduction: roughly 48,000 words organised by subsystem, with the reasoning and sources behind every choice. They are meant to be audited against and referred to, not read front to back.

Reading order#

  1. Vision and north star
  2. Scope and realism contract
  3. Research landscape
  4. World model and scenario
  5. Agents, cognition, and memory
  6. Society, economy, governance, and conflict
  7. Architecture and data
  8. Time, emergence, and observation
  9. Models, cost, operations, and security
  10. Validation and experiments
  11. Roadmap and gates
  12. Risks, decisions, and open questions
  13. Independent audit guide
  14. Sources
  15. Determinism, replay, and state integrity
  16. Parameter registry
  17. Cost model and model selection
  18. Initial conditions and authorship
  19. Lab charter and research program
  20. Experiment card template and R0 continuity protocol

Documents 15–18 were added by the first audit pass. Documents 19–20 and ADR-009 were added by pass 2 to connect the terrarium to a falsifiable lab program, define the first experiment, and prevent operational identity from being confused with consciousness or human survival.

Architectural decisions are in wiki/adr/. The reviewer can copy the audit report template to AUDIT-REPORT.md.

Non-negotiable principles#

  • The simulator owns truth. Language models propose intentions and actions; typed systems decide what is possible and what occurs.
  • No invisible storyteller. Exogenous events arise from explicit, state-conditioned hazard models. Narrative summaries never mutate the world.
  • Conservation before conversation. Food, water, fuel, materials, labor, distance, health, and time cannot be invented in prose.
  • Private viewpoints. Agents see local observations, testimony, records they can access, and fallible memories—not global state.
  • Consequences persist. Death, injury, burned buildings, broken trust, law, debt, and ecological depletion survive context-window resets.
  • History is replayable. All mutations are events with causal parents, seeds, actor, authorization, and before/after hashes.
  • Silence is allowed. Agents do not need to manufacture drama. Uneventful seasons are valid.
  • Cost is a world constraint. Budgets degrade cognition gracefully; they never corrupt physics or silently stop time.
  • Extinction is allowed. It is a terminal historical outcome, followed by preservation and optional counterfactual forks—not a forced reset.
  • Claims climb one level at a time. Behavioral continuity is not consciousness; human-model fidelity is not human continuation.
  • No human persona intake yet. Human-subject work requires a separate ethics, consent, privacy, retention, and withdrawal protocol.
  • No implementation before audit. Phase 0 ends only after all blocking audit findings are resolved or explicitly accepted.

What “done with planning” means#

Planning is complete only when the audit can trace each promised behavior to:

  1. an authoritative state variable;
  2. a transition rule or approved LLM decision boundary;
  3. observable evidence;
  4. an invariant or validation test;
  5. a failure policy;
  6. a phased implementation gate.

The current document is intended to reach that standard conceptually. Numerical parameters remain provisional until Phase 1 calibration; provisional values must be labeled and stored with provenance.

Communication briefLANDING-PAGE-BRIEF.md
739 words
Inside this chapter

December landing page#

The simple story#

We are trying to build a little world that keeps living while no one is watching.

Eighteen artificial residents begin in a valley. They remember kindness, hold grudges, build homes, make rules, fall in love, grow old, and live with what they have done. Leave for two days and their lives continue without you.

The farthest dream is that learning what holds an artificial life together might someday illuminate what lets a human life remain itself through change—and whether something essential could ever continue beyond the body.

That is the hope. It is not a claim that we have solved consciousness or immortality.

Hero copy#

Eyebrow: A LONG-TERM DREAM, BUILT IN THE OPEN

Headline:

A little world that keeps living.

Body:

Eighteen artificial residents begin in a valley. They remember kindness, hold grudges, build homes, make rules, fall in love, grow old, and keep living—even while no one is watching.

Primary action: Enter the dream

Secondary action: Read the plan

Status: The world is being designed now. Nothing here is a claim of conscious AI.

Page flow#

1. The experience#

Leave for two days. Return to lives that continued without you.

You might return to a wedding, a new law, a burned granary, a bitter election, a migration, a war—or an empty valley. Nothing happens because a storyteller wanted a dramatic scene.

2. The beginning#

Build a world. Let lives unfold.

The current design cohort—eighteen artificial residents—begins in a small valley. They must find food, form relationships, make rules, remember the past, and live with the results of their choices. The final cohort size remains an experimental decision.

We create the people and the boundaries, but never choose their government, war, religion, survival, or ending. The authored assumptions remain visible.

3. Why a whole world?#

A self is made with others.

People are shaped by bodies, memories, relationships, work, scarcity, culture, loss, and time. A resident should be more than a clever voice in a chat window.

4. What the dream must be able to hold#

  • Ordinary days as well as extraordinary ones.
  • People who want different things—including power, safety, love, revenge, duty, glory, peace, or sacrifice.
  • Things residents devise themselves: projects, customs, alliances, offices, elections, rebellions, migrations, and new ways of living.
  • Consequences that persist: trust, injury, debt, buildings, law, ecological loss, death, and extinction.

5. The distant hope#

Could a life continue beyond the body?

We do not know. December cannot answer that today, and a convincing artificial resident would not prove consciousness or human survival. The question is a compass, not a promise.

6. Built in the open#

The code, assumptions, failures, costs, and histories will be open to inspection. Engineers, researchers, artists, world-builders, and skeptics are invited to help.

Final action: Bring us a possibility

Code-native hero ASCII#

                         a life, observed

                              ·
                         ─────●─────
                       ╱      │      ╲
                  memory   choice   others
                       ╲      │      ╱
                         ─────●─────
                              │
                         time continues

                    WHAT MAKES YOU, YOU?

World-to-person diagram#

body ─────┐
memory ───┤
others ───┼──> a life through time ──> identity
world ────┤
choice ───┘

Language rules#

  • Use short sentences and ordinary words.
  • Lead with the felt experience of returning to a world that continued without you.
  • Say "artificial people" before "agents."
  • Say "small world" before "causal kernel."
  • Treat consciousness as a question.
  • Treat immortality as a distant hope, never a product promise.
  • Never describe operational identity scores as evidence of consciousness or human survival.
  • Let the page feel like an invitation into a dream. Keep research language below the emotional story and use it only to establish honesty.
  • Keep Wega Labs clearly identified as the company and AI lab. The distinction is: Wega is the lab; December is one of its long-term dreams and research projects.
  • Do not solicit human life histories or nominations. Human-participant work requires a separate first-party, private, reviewed protocol.
  • Keep technical proof available below the main story for researchers and contributors.

Implementation status#

The 2026-08-01 copy pass completed the central corrections in both implementations:

  1. The human-participant cards, modal, form handling, and public-life-history request were removed.
  2. Calls to action now ask people to follow the build, ask a question, or bring a possibility.
  3. Hero, sections, metadata, sharing copy, and schema now lead with the dream rather than consciousness research.

Still required before the wider Wega identity is considered settled:

  1. Rename or clearly separate “December Sato,” the unrelated autonomous Mac mini agent.
  2. Publish a stable claims-and-limitations page when the research documentation has a public canonical URL.
Part II

The world specification

The complete product, world, agent, society, architecture, validation, and roadmap specification.

Normative designwiki/00-vision-and-north-star.md
1,113 words
Inside this chapter

00 — Vision and North Star#

One-sentence vision#

Build a tiny, persistent civilization that produces unscripted but causally explainable history while remaining cheap enough to run continuously and transparent enough to audit like an experiment.

For the larger Wega Labs thesis, this world is an instrument for studying persistent synthetic agency—not a direct test of consciousness or immortality. The research claims ladder and first falsifiable experiments are defined in 18.

Terrarium versus autonomous-agent society#

The phrases describe different design commitments.

Dimension “Society of autonomous agents” “Observable AI terrarium”
Center of attention Agent personalities and autonomy The coupled system: people, material world, institutions, ecology, and observer instruments
Typical implementation Several LLMs chat and call tools Deterministic/stochastic simulation plus bounded LLM decisions
Source of events Often dialogue, prompts, or a narrator State transitions and conditional hazards
Truth Frequently distributed through prose One authoritative typed world state
Observer Watches transcripts Can inspect causal chains, maps, metrics, beliefs, and replay
Surprise Model improvisation Emergence from interacting mechanisms, with language adding interpretation
Failure mode Expensive role-play; incoherent resources and memory Over-engineered toy ecology or legible but less theatrical early runs
Scientific posture Demonstration of agents Instrumented computational world with declared validity limits

We still want a society of autonomous agents, but it lives inside a terrarium. “Terrarium” reminds us that habitat, constraints, observation, and experimental control matter as much as minds.

The desired return-after-two-days experience#

The home screen should answer four questions in under a minute:

  1. What changed? Population, stores, territory, buildings, offices, law, factions, health, ecological indicators.
  2. What happened? A ranked timeline of births, deaths, projects, votes, disputes, discoveries, migrations, disasters, and battles.
  3. Why? A causal graph from outcome back to decisions, observations, physical constraints, and random draws.
  4. What might have happened? Optional shadow branches from selected decision points.

An illustrative—not scripted—history might read:

Year 3, late dry season: low winter rainfall cut the east stream to 38% of its median flow. The council rationed irrigation, but two households concealed grain. A child’s testimony shifted the election. The defeated steward refused the audit, left with four supporters, and occupied the upstream weir. A confrontation injured three people. Close-contact care spread a respiratory infection. With labor unavailable, the millet harvest failed; eight residents migrated and the original settlement ended the winter with five survivors.

Each sentence must link to underlying events and evidence. The system is a failure if it can produce the paragraph but not justify it.

Experience pillars#

1. Causal depth#

Consequences should cross systems. Rain affects soil and streams; those affect labor and harvest; stores affect bargaining power; institutions affect distribution; distribution affects health and loyalty; injury affects productive capacity. Modules cannot be isolated minigames.

2. Persistent personhood#

Agents have bodies, histories, relationships, obligations, skills, incomplete knowledge, and limited time. Their choices should reflect these without collapsing into fixed personality stereotypes.

3. Institutional emergence#

There is no compulsory “democracy module.” Residents can create offices, voting rules, councils, customary law, ownership norms, sanctions, alliances, or authoritarian arrangements using a bounded institutional grammar. Elections become possible when people enact them.

4. Constructive agency#

Residents can propose genuinely new combinations of known capabilities: buildings, work processes, symbols, schedules, organizations, records, and policies. New things enter reality through blueprints, bills of materials, labor, tests, and use—not by being mentioned.

5. Legible surprise#

Surprise is valuable only if it is neither pre-authored nor inexplicable. We optimize for retrospective inevitability: after seeing the causes, the outcome makes sense, though it was not obvious beforehand.

6. Long-horizon durability#

The world survives restarts, provider outages, malformed model output, budget limits, and schema upgrades. It can run headlessly for weeks. Every canonical history is versioned and backed up.

Success criteria#

The project is successful when all are true. Every criterion below that uses a label — "governance form," "faction," "cascade" — requires a pre-registered operational definition before the runs that test it, per 17. A criterion whose definition is chosen after seeing the results is not a criterion.

  • In blinded review, observers distinguish independent causal histories and can correctly identify major causes from evidence.
  • Replays reproduce identical state hashes at every checkpoint within a pinned environment — same code commit, config, dependency lockfile, interpreter build, CPU architecture, and model-response cache. Replay hard-fails rather than calling a live model on a cache miss. Portable bit-identity across machines is not claimed unless the kernel adopts integer-valued state; see 14.
  • Removing any one major pressure—seasonality, resource conservation, social memory, institutions—measurably changes population-level outcomes relative to a null demographic model, so that small-population noise is not mistaken for mechanism.
  • Agents accomplish multi-day construction and institutional projects without developer-authored sequences.
  • At least three governance forms arise across seeded runs that are structurally distinct under the pre-registered definition and persist beyond a declared duration — including in runs seeded without an inherited charter.
  • Rare cascades such as settlement fission, introduced-disease outbreaks, war, or extinction occur in some parameter regions but are not guaranteed. At this population size, imported acute infections should show substantial stochastic fade-out and occasional large outbreaks across ensembles; intermediate outbreaks remain possible (15 §D).
  • An auditor can trace every material mutation to code version, command, preconditions, authorization, and random seed.
  • A seven-day soak test completes without human repair, unbounded costs, state corruption, or narrative/world divergence.
  • Observed cost per simulated day stays within the owner-approved envelope (16), and model refusal rates are measured and reported alongside any behavioral finding.

Explicit anti-goals#

  • A scientific forecast of real human societies.
  • A conscious-being claim or moral-patient experiment.
  • A chatroom with biographies and a simulated clock.
  • A god-game where the observer secretly nudges events in the canonical run.
  • Maximum population, graphical fidelity, or token consumption.
  • A system in which agents write and execute unrestricted code.
  • A world optimized to produce violence or catastrophe on demand.
  • A single giant prompt containing the entire civilization.

Lab-level success#

The canonical valley is a compelling exhibit and operational test, not by itself a scientific cohort. Wega Labs earns the “lab” label by producing registered experiments, baselines, reconstructable artifacts, negative results, and changes of belief. In the first year, success means a credible longitudinal identity benchmark, controlled model-transplant and social-grounding results, and a replayable causal-agent testbed—even if the canonical world never produces a spectacular war or extinction.

Working title and terminology#

“December” is the project codename. “The Living Terrarium” is the product concept. “Canonical world” means the one continuous history visible to the observer. “Shadow world” means a non-canonical fork used for analysis. “Resident” means an individually simulated person. “World kernel” means the authoritative transition system. “Director” means a read-only summarizer and camera planner; it has no mutation authority.

Normative designwiki/01-scope-and-realism-contract.md
1,649 words
Inside this chapter

01 — Scope and Realism Contract#

Why an early agrarian settlement#

Modern society would require finance, mass media, electricity, industrial supply chains, bureaucracy, transport, medicine, law, and thousands of invisible institutions before one convincing day could occur. A prehistoric hunter-gatherer band is tractable but offers less construction, durable surplus, office, property, and formal governance.

The recommended starting point is a fictional early agrarian frontier, inspired by but not presented as any specific culture. It has mixed farming, gathering, hunting, fishing, pottery, fiber, wood and stone construction, storage, simple water works, and oral/marked records. This gives us:

  • seasons, harvests, spoilage, soil decline, and food storage;
  • households, kinship, commons, inheritance, and inequality;
  • multi-day buildings and infrastructure;
  • enough surplus for specialists and offices;
  • nearby groups, trade, migration, territorial conflict, and epidemic contact;
  • few enough technologies that recipes and capabilities can be enumerated.

The fictional setting avoids claiming historical authenticity. Every parameter inspired by scholarship must carry provenance, confidence, and sensitivity range.

Initial scope#

Current design cohort—not a validated optimum#

  • 12 full-cognition adult residents in 6 households.
  • 4 child dependants and 2 elders initially use lightweight policy cognition but retain bodies, relationships, needs, memory, and promotion to full cognition when relevant.
  • One neighboring band begins as an aggregate population with stores, territory, disposition, and internal factions. Named individuals are materialized when contact, migration, diplomacy, or conflict makes them relevant.
  • Births create lightweight persons; maturation promotes them. Death is permanent.

This hybrid “level of detail” policy spends tokens on consequential perspectives without turning dependants or outsiders into mere counters.

The candidate population is marginal, and that is a design constraint rather than a caveat#

Eighteen people is below many modeled recommendations for a viable isolated human population, while historical small-founder cases depended on conditions and outside contact that are not clean substitutes for December (15 §C). These sources bound questions; they do not establish a universal threshold or validate this cohort.

So the honest statement is that December's founding group is marginal, not impossible — and the interesting design question is what makes the difference. Four consequences follow, and all four are binding:

  1. Demographic stochasticity dominates. With expected births and deaths each in the single digits per decade, run-to-run variance is comparable to the mean. A substantial fraction of runs will end for no interesting reason. Validation must therefore compare every claimed collapse cascade against a null demographic model with no social mechanisms, or it cannot distinguish causality from noise.
  2. Partner availability matters immediately. Chance age/sex-ratio skews and socially permitted matching can constrain births before genetic load becomes the dominant concern. December therefore models partner eligibility explicitly and treats any literature-derived multiplier as a sensitivity range, not a law.
  3. Network severance is one plausible failure mechanism. Historical cases suggest outside exchange and migration can matter, but they do not identify one universal cause of small-settlement failure. December tests boundary connectivity as a mechanism instead of importing a historical narrative.
  4. In-migration and exogamy are therefore load-bearing mechanics. The boundary world is not primarily a source of conflict and disease; it is the lifeline. It must exist from Phase 3, not Phase 5, or every institutional experiment runs inside a closed, guaranteed-declining population.

One neighbor of similar size is not enough — two settlements of 18 give a pool of ~36, roughly a quarter of the minimum viable mating network of ~150–500. See 03 for the recommended regional-aggregation mechanic.

If the observer experience depends on generational turnover or inherited institutions, the founding size must be revisited. Gate 1's ensembles report extinction rate as a function of founding population, and that curve sets the final number.

Map#

A hex or square grid representing roughly a valley-scale walkable region. Cells store elevation, slope, soil class, moisture, surface water, vegetation/fuel, resource stocks, ownership/use claims, structures, and current occupants. Exact scale is a Phase 1 decision; distance and travel time must be internally consistent.

Technology boundary#

Initial capabilities include fire, shelters, storage pits, wood/stone tools, cordage, baskets, pottery, fishing, hunting, gathering, two staple crops, and simple irrigation. Metal, writing, animal traction, wheeled transport, and complex fortification are absent until deliberately added as tested expansion packs.

Time horizon#

  • Target canonical pace: configurable, initially 1 real hour = 1 simulated day while attended and up to 1 real hour = 3 simulated days while unattended.
  • Kernel tick: event-driven with daily physiological/ecological boundaries.
  • Cognition cadence: on meaningful triggers, not every tick.
  • Initial soak target: seven real days / several simulated months, then multi-year accelerated experiments.

The mapping is not sacred. We choose it only after cost and behavioral tests. Long-term historical change needs accelerated experimental branches even if the canonical world moves slowly.

The realism contract#

“Super realistic” is otherwise an invitation to hide hand-waving. This project uses six testable meanings.

R1. Material realism#

Resources have quantities, locations, quality, ownership/custody, and transformations. Actions consume time, energy, tools, inputs, and access. Outputs cannot exceed inputs plus modeled growth or extraction.

Required evidence: stock-flow ledgers, conservation/property tests, capacity limits, spoilage and loss events.

R2. Temporal and spatial realism#

People cannot be in two places, learn news instantly, complete work without elapsed time, or respond before observing an event. Travel, communication, construction, disease, crops, injuries, and institutions operate at different timescales.

Required evidence: interval overlap checks, path/travel events, message provenance, scheduled processes.

R3. Epistemic realism#

World truth, an agent’s observations, their beliefs, public records, rumor, and retrospective narration are separate objects. Memory may be lossy or biased, but the event log is not.

Required evidence: information lineage for every decision; tests preventing leakage of hidden state.

R4. Behavioral plausibility#

LLMs choose among feasible actions using needs, commitments, norms, goals, emotions-as-appraisals, relationships, and bounded plans. Behavior is neither perfectly rational nor unconstrained improvisation. Traits modulate decisions; they do not dictate them.

Required evidence: scenario tests, repeated-seed behavioral distributions, ablations, contradiction rates, independent human review.

R5. Structural realism#

Macro events must arise from plausible micro mechanisms and feedback loops. “War” is mobilization, logistics, communication, risk, combat encounters, injury, morale, and political aftermath—not a scalar event. “Election” is eligibility, candidacy, information, voting, counting, disputes, transition, and powers.

Required evidence: causal graphs and module-level ODD descriptions.

R6. Empirical humility#

The model declares what is calibrated, borrowed, speculative, or designed. It reports sensitivity and uncertainty rather than presenting a cinematic outcome as human science.

Required evidence: parameter registry (15), provenance, ensemble runs, validation report, “known invalid” list.

R7. Scale honesty#

A mechanism must be valid at December's population size, not merely valid somewhere. Parameters and model structures borrowed from populations of thousands or millions are frequently meaningless at n=18, and a mechanism tuned until it "works" at this scale may be encoding a rate that cannot physically exist.

The disease module is the worked example: no acute, directly transmitted, immunizing infection can be endemic in eighteen people, because critical community size for such infections is in the hundreds of thousands (15 §D). An SEIR model tuned until epidemics appear at a satisfying frequency would be fabricated epidemiology wearing the costume of rigor.

Required evidence: for every subsystem, an explicit statement of the population range over which its structure is valid, and a test that the subsystem behaves correctly — including degenerately, where that is the correct behavior — at n=18. Where the honest answer is "this mechanism cannot operate at our scale," the subsystem is redesigned around one that can, not tuned until it produces output.

Realism tiers#

Every subsystem receives a tier so ambition does not become accidental scope.

Tier Meaning Example
0 Placeholder Interface only; not allowed in canonical long runs Combat returns a fixed result
1 Causally coherent Correct stocks, timing, preconditions, broad feedback Crop yield responds to labor, water, soil, weather
2 Pattern-calibrated Reproduces several target qualitative/quantitative patterns Seasonal hunger and storage reduce variance in intake
3 Domain-reviewed Assumptions and results reviewed by a relevant expert Epidemic module reviewed by an epidemiologist

Canonical launch requires Tier 2 for food/water, energy/health, demographics, construction, information, and core social exchange; Tier 1 for conflict and government; no Tier 0 modules enabled. Tier 3 is aspirational.

Laws of the world#

These are invariants, not configurable personality choices:

  1. Identity: every entity and event has a stable ID.
  2. No negative stocks: transfers and consumption are atomic and bounded.
  3. Conservation: mass-like resources are created only by declared source processes and destroyed only by declared sinks.
  4. Single location: a person occupies one location or one travel edge at a time.
  5. No retrocausality: decisions use information available before their timestamp.
  6. Mortality: dead agents cannot act; estates and obligations transition by a declared rule.
  7. Authority: an institutional action requires a currently valid capability.
  8. Idempotency: repeated command delivery cannot duplicate its effect.
  9. Seeded randomness: all stochastic draws use named streams recorded with events.
  10. Narrative impotence: summaries, cameras, and observer queries have no world-write capability.

What can be random#

Randomness represents modeled uncertainty or variation, not authorial convenience. Each hazard has:

  • eligibility conditions;
  • a rate or probability conditioned on state;
  • a named RNG stream and draw ID;
  • severity distribution;
  • spatial footprint;
  • downstream mechanics;
  • provenance/confidence;
  • tests for impossible and degenerate rates.

Examples include rainfall, lightning ignition, conception, accident, pathogen introduction, transmission, hunting success, and combat injury. A generic “dramatic event” roll is forbidden in the canonical world.

Ethical and interpretive limits#

  • The agents are software personas; we avoid claims of consciousness, suffering, or moral equivalence to people.
  • Population attributes should not encode race, ethnicity, or real-world protected-group stereotypes.
  • Reproduction and kinship remain non-explicit, abstract state transitions with consent constraints.
  • Violence is represented analytically, without graphic content and without optimizing for cruelty.
  • The observer is told when a claim is a model inference, a resident belief, or direct world state.
  • No outcome should be used to make policy claims about real communities without an entirely separate, domain-reviewed validation effort.
Normative designwiki/02-research-landscape.md
1,523 words
Inside this chapter

02 — Research Landscape and Starting Points#

Decision in brief#

No single open-source project is the base we need. The strongest approach is a composed architecture:

  • a new typed causal kernel, likely built with Mesa’s Python ABM/event primitives;
  • Agentopia-inspired resident cognition, contact, review, memory, checkpoints, and long-horizon operation;
  • GovSim and Melting Pot as libraries of social dilemmas and evaluation scenarios;
  • VillagerAgent/APT ideas for decomposing construction into dependency graphs;
  • disease mechanics adapted from established individual-based epidemic models, simplified for a tiny population;
  • an optional Craftium/Luanti or custom 2D adapter only after the headless world works;
  • LiteLLM for provider routing and Langfuse/OpenTelemetry-style tracing.

“Inspired/adapted” does not mean copy-paste. License compatibility and actual code quality must be checked before code reuse.

Candidate matrix#

Project What it gives us What it does not give us Proposed use
Agentopia Long-running social simulation; Plan→Contact→Activity→Review; self-managed memory; append-only records/checkpoints; published token/cost figures A grounded material world; mature ecosystem; a LICENSE file — the repo is three commits deep and legally unlicensed despite a README claim of MIT Cognitive architecture and cost reference; do not vendor code
Stanford Generative Agents Memory stream, reflection, planning, believable daily behavior Resource conservation, institutions, long-horizon operations Foundational cognition patterns and evaluation ideas
AI Town Runnable multiplayer browser town and agent conversations Deep material, demographic, ecological, or political causality UI/product reference only
MiroFish Multi-agent simulation and report generation around supplied seed material Persistent embodied causal world; AGPL-3.0, whose obligations trigger on network use — the one genuine licensing trap in the candidate set Observer-report inspiration only; do not vendor
Concordia Configurable generative-agent simulations with a Game Master and components Game Master can become an omniscient narrative bottleneck Experimental harness; borrow component separation, constrain GM authority
Mindcraft Mature Minecraft/Mineflayer action surface; OpenRouter; gathering/crafting/building; multi-agent support Minecraft is game-balanced, not historically/ecologically realistic; code execution is dangerous Fast embodied prototype or later adapter; arbitrary code disabled
Craftium + Luanti Fully open sandbox; Gymnasium/PettingZoo APIs; synchronous stepping explicitly for slow agents such as LLMs; ICML 2025 paper LGPL-2.1-or-later plus CC BY-SA media, not permissive; development slowed since Feb 2026; significant integration work Preferred 3D research adapter candidate after Phase 4 — and a licensing decision, not just a technical one
MineLand Multi-agent Minecraft benchmark and tasks Benchmark rather than persistent civilization Evaluation/task ideas
VillagerAgent Collective task decomposition and dependency-aware execution Full society/world simulation Project compiler and work-allocation patterns
APT Converts text construction goals to structured blueprints Material society and institutions Blueprint compilation inspiration
GovSim (Piatti et al., NeurIPS 2024) Commons-governance scenarios with measurable outcomes; MIT Persistent bodies, ecology, construction, life history; frozen since Jan 2025 Governance tests and institutional mechanism library
GovSimElect Elected/fixed/leaderless leader comparison over GovSim's fishery It is a five-star personal fork of GovSim, not an upstream project — earlier drafts miscited it Cite as a fork; useful only for the election variant
Agent Ballot Box Commons framework with a voting/leadership component Whole polity/world; zero adoption, dormant since mid-2025 Weak citation; retained for completeness only
Melting Pot 50+ social substrates and 256+ scenarios covering cooperation, competition, deception, reciprocity, trust, coalition behavior Persistent historical world and LLM-native residents Regression scenarios and social-mechanism catalog
MoralAgentSim Prehistoric agents that hunt, gather, share, communicate, reproduce, and fight; ACL 2026 Main (oral) Studies moral evolution, not a persistent material world; unknown engineering maturity The closest published analogue to December's premise — highest-value comparative spike; borrow no claims uncritically
AgModel Open forager–farmer transition model with annual demography/environment and event-based subsistence LLM cognition and rich institutions Parameter/mechanism reference for calories, labor, stores, birth/death
Artificial Anasazi Household settlement, land productivity, drought and food surplus Its own page warns it is unverified; ecology alone does not explain observed collapse Warning and test case, never historical ground truth
Simulating Forager Mobility / archaeology ABMs Mobility, resource distribution, water tethering, population dynamics, fission–fusion Production software architecture Mechanism and validation-pattern references
Covasim / Starsim / OpenABM Individual disease states, contact networks, stochastic transmission, testing discipline Tiny premodern generic disease out of the box Adapt architecture and validation, not COVID parameters
WarAgent Structured LLM diplomacy/war simulation World-war abstraction; little local logistics/embodiment Diplomacy action grammar reference only
0 A.D. Historical RTS engine with economy, building, units, battle Designed balance, large codebase, wrong scale/cognition Art/interaction inspiration; reject as authoritative base
Widelands Detailed worker/ware/building economy; active open-source project Economy-game assumptions; C++ integration cost Recipe/logistics reference; possible visualization research
Unknown Horizons Economy and settlement UI ideas Dormant — last release Jan 2019; the Godot rewrite is a separate repo with no playable content; "Python/Godot" conflates two codebases UI inspiration only; do not plan around it
Freeciv Mature client/server and rulesets Civilization-level turns, not individual lives Ruleset/server architecture reference only
OPA Auditable policy-as-code and capability checks Natural social process or flexible emergent constitution Later use for infrastructure authorization, not residents’ full law
LiteLLM Unified gateway, routing, budgets, retries across providers Simulation semantics Model gateway candidate
Langfuse LLM traces, cost, latency, prompt/version observability Canonical world events Operations telemetry linked to event IDs

Why not start by forking an RTS#

0 A.D., Widelands, Freeciv, and Unknown Horizons are impressive, but their rules encode gameplay abstractions: accelerated production, balance-driven yields, omniscient players, fungible units, and combat designed for control. Retrofitting individual knowledge, kinship, physiology, continuous life history, institutional authority, and scientific event provenance would fight the engine.

We should still mine them for:

  • production-chain and recipe data structures;
  • pathfinding and job reservation;
  • building placement and visualization;
  • client/server separation;
  • deterministic replay lessons;
  • moddable rulesets and content pipelines.

Why Mesa is the leading kernel scaffold#

Mesa is an Apache-2.0 Python agent-based modeling framework with agent/model primitives, spatial components, data collection, browser visualization, and support for hybrid step/event scheduling. It aligns with the need for scientific experiments and fast iteration.

Status as of August 2026 (verified during the audit pass — earlier drafts of this wiki were out of date):

  • Current stable is 3.5.1 (March 2026), with a 4.0 line in alpha.
  • Discrete-event and hybrid scheduling are stable, not experimental: model.schedule_event(), model.schedule_recurring(), run_for(), run_until().
  • The mesa.experimental.devs simulator classes are deprecated since 3.5.0 and removed in 4.0. Any design targeting them is already stale.
  • mesa.space is maintenance-only, and SolaraViz carries breaking-change risk across minor releases.

That last pair is the important finding: the parts of Mesa December would lean on hardest are the parts under active churn or reduced maintenance. It strengthens rather than weakens the case for the narrow interface. Regardless of the choice:

  • event sourcing, deterministic RNG streams, typed commands, and causal provenance must be ours;
  • performance must be benchmarked before commitment;
  • the core domain must not inherit framework-specific serialization everywhere;
  • a narrow SimulationRuntime interface must allow replacing Mesa without touching domain code;
  • the determinism constraints in 14 §D2 apply to any runtime — a framework that iterates unordered collections or reads wall-clock time inside the step loop is disqualified regardless of its other merits.

The decision gate is a Phase 1 spike comparing Mesa against a small custom event loop on determinism, throughput, profiling, checkpointing, and developer clarity. Given that December needs its own scheduler semantics, its own event store, its own RNG discipline, and its own provenance model, the honest prior is that Mesa earns its place as a spatial and data-collection toolkit rather than as the kernel.

Research methods we adopt#

ODD#

The Overview–Design concepts–Details protocol is a standard way to describe individual/agent-based models. Each kernel subsystem will have an ODD-compatible specification covering purpose, entities/state, process/scheduling, design concepts, initialization, input data, and submodels.

TRACE-style evaluation#

Documentation must include problem formulation, conceptual model, implementation verification, parameterization, data evaluation, model analysis, and output corroboration. We use this as a discipline even though December is not initially a scientific policy model.

Pattern-oriented validation#

We should test multiple patterns simultaneously rather than tune one headline outcome. For example, a food system should reproduce plausible seasonality, labor bottlenecks, storage buffering, and starvation response—not merely an average annual yield.

Ensemble and ablation experiments#

One beautiful run proves almost nothing. Each version must run many seeds, sweep uncertain parameters, and remove proposed mechanisms to demonstrate which outcomes depend on which mechanisms.

Rejected shortcuts#

  • LLM as environment simulator: fluent but non-conservative, difficult to reproduce, and vulnerable to prompt drift.
  • A “chaos slider”: makes drama authorial rather than emergent.
  • All agents awake every minute: cost-heavy and socially noisy.
  • Infinite crafting via text: breaks the material contract.
  • One shared memory database: leaks private knowledge.
  • One model provider: fragile and wastes heterogeneous token balances.
  • World state only in vector search: approximate retrieval cannot be authoritative.
  • Unrestricted shell/code tools: unnecessary security and integrity risk.
  • Trusting a README's license claim, or a repository that turns out to be a fork. The audit pass found both errors in this wiki's own bibliography. Verify the LICENSE file and the upstream before any reuse decision.
  • Borrowing a mechanism that is invalid at our population size. A model validated on thousands of agents may be meaningless at eighteen — see realism contract R7 in 01.
Normative designwiki/03-world-model-and-scenario.md
2,338 words
Inside this chapter

03 — World Model and Founding Valley Scenario#

Scenario: Founding Valley#

A flood-fed tributary crosses a temperate valley with upland woodland, marsh, grassland, arable patches, clay, stone, fish, and migratory prey. Twelve adults from several households have established a new settlement after leaving an older community. They share a language and inherited customs.

The founding state is declared, randomized, and varied—not unauthored — see 17, which is binding on this section. Earlier drafts described the founders as disagreeing "about property, leadership, risk, and relations with a neighboring mobile band," and described land claims as "ambiguous" and assembly procedure as "unsettled." Those were not neutral initial conditions; they were the plot, pre-installed. If the settlement later fractures over property and elects a leader to resolve it, we would have observed only what we planted. Generators do not remove authorship: their distributions, constraints, correlations, exclusions, and accepted seeds are all research decisions.

What the scenario may specify is material and structural:

  • household composition, ages, kin links, and dependency ratios sampled from declared demographic ranges;
  • per-household stores, tools, and skills, unequal because households differ in size and history;
  • actual land claims with actual evidence and actual overlaps — ambiguity is then a derived property of the claim ledger, not an authored mood;
  • distances from each dwelling to water, arable land, and fuel;
  • a partly planted first crop, an unfinished irrigation ditch, and shelters of differing quality, each with real material state;
  • values sampled per resident from the declared vocabulary in 04, with seed and distribution recorded.

The founding charter carries three inherited rules — violence is forbidden within the central hearth area, the communal grain store requires two key-holders, and disputes may be called before an assembly. These are legitimate as inherited culture from the community the founders left, with recorded provenance. They are not a designed constitutional crisis: either the charter specifies the assembly's procedure or there is no assembly. Deliberately underspecifying a rule so it can be fought over is authorship.

A no-charter control arm is mandatory in every institutional ensemble. If governance emerges only when a charter is seeded, then the charter is producing the institutions and the emergence claim is void.

"Neither utopian nor already doomed" is a design intent, and it becomes a specification only when measured: initial stores cover a stated number of days at current consumption, and the ensemble's one-year survival rate under scripted policies falls within a declared band.

Authoritative state domains#

Geography and environment#

Each cell or region tracks:

  • coordinates, elevation, slope, surface type, traversability;
  • soil texture, fertility pools, erosion, moisture, and cultivation state;
  • surface water flow/volume and groundwater proxy;
  • vegetation biomass by functional group and burnable fuel;
  • wild food and prey stocks with regeneration/mobility;
  • weather exposure, temperature proxy, precipitation, wind, and fire state;
  • constructed features, paths, fields, claims, and occupancy.

Weather is generated from a seasonal stochastic process with autocorrelation and scenario-level climate parameters. We do not need a meteorological model; we do need wet/dry persistence, seasonal temperatures, correlated spatial effects, and extreme tails. Weather feeds hydrology, crops, vegetation, fire, travel, and health.

People and bodies#

Every person has:

  • age/life stage, household/kin links, location and activity;
  • energy reserve, hydration, temperature/exposure, fatigue, sleep debt;
  • disease states, injuries, functional limitations, pregnancy/infant state where applicable;
  • skills, practiced proficiency, known techniques, and teaching links;
  • carried inventory, access rights, claims, debts, obligations, and offices;
  • cognition tier and next activation conditions.

Sex/gender and reproduction require a later design note and review. The kernel only needs the minimal attributes required for demographic transitions and consensual household choices; it must not infer behavior from stereotypes.

Things and resources#

Resources are typed lots, not prose:

lot = {
  kind, quantity, unit, quality, condition,
  location_or_container, custodian, claimants,
  produced_at, expires_at?, provenance_event
}

Initial kinds include potable/non-potable water, edible plants, grain, fish/meat, seed grain, wood by size, fiber, clay, stone, hides, fuel, tools, vessels, medicine proxies, and waste. Quality affects nutrition, durability, spoilage, and disease risk.

Structures and projects#

A structure is a spatial assembly of components with condition, capacity, access, function, and maintenance needs. A project holds:

  • proposed design and purpose;
  • site and access authorization;
  • dependency DAG;
  • bill of materials;
  • labor tasks and skill/tool requirements;
  • work completed and defects;
  • owner/steward and future maintenance schedule.

There are no instant buildings.

Relationships and groups#

The social graph distinguishes:

  • kin/household relationship;
  • familiarity and frequency of contact;
  • trust by domain, affection, fear, grievance, respect, perceived competence;
  • obligations, gifts, debts, promises, testimony, and conflicts;
  • groups, membership, roles, entry/exit rules, shared assets, and declared purpose.

Relationships are directional and evidence-based. “Trust” is not one universal score: a resident may trust someone’s farming judgment but not their honesty with stores.

Institutions and public artifacts#

Institutions are persistent rule bundles plus roles and assets. Public artifacts include enacted rules, proposals, vote records, office grants, judgments, boundary markers, tallies, calendars, and agreements. Literacy is unnecessary: records can represent witnessed oral commitments, tokens, marked clay, cords, or maintained memory, with durability and accessibility properties.

Core physical processes#

Food and energy#

The kernel models energy in consistent abstract units, with documented conversion to calories only when calibrated. Daily needs depend on age/life stage, condition, temperature, and exertion. Food lots carry energy/nutrient proxy, water content, contamination, and spoilage state.

Food pathways:

  • gathering depletes renewable patches with diminishing returns;
  • hunting/fishing success depends on stock, location, skill, tools, effort, and stochastic encounter/capture;
  • crops pass through phenological stages conditioned on planting window, moisture, temperature, soil, labor, pests, and damage;
  • processing changes edibility, storage life, and labor cost;
  • storage capacity, pests, moisture, theft, and spoilage matter;
  • seed consumption trades present survival against future yield.

Water#

Water exists in sources and containers. Collection takes time; treatment methods reduce hazards; irrigation diverts finite flow; upstream activity affects downstream access and quality. Wells and aqueducts are future expansions. This system is a natural source of both cooperation and conflict.

Materials, tools, and maintenance#

Extraction changes the environment. Tools have condition and task multipliers; broken tools become repairable components or waste. Buildings decay from exposure, use, pests, fire, and deferred maintenance. A settlement can grow itself into a maintenance trap.

Health, injury, and disease#

Health is not a hit-point bar. The first release uses:

  • nutritional reserve and chronic undernutrition;
  • dehydration and exposure;
  • fatigue/sleep;
  • task/accident injuries by body-function category;
  • wounds, contamination, healing, impairment, and mortality risk;
  • the five-mechanism disease model below.

Disease must be restructured for a population of eighteen#

A single generic SEIR pathogen — the earlier design — is the wrong model at this scale. Critical community size for acute, directly transmitted, immunizing infections is in the hundreds of thousands; December's population is four orders of magnitude below it (15 §D). Such a pathogen will either fade out immediately and contribute nothing, or, if tuned until epidemics appear at a satisfying rate, encode a rate that cannot exist. The second outcome is risk R-07 wearing epidemiological costume.

Five mechanisms with genuinely different persistence logic replace it:

Mechanism Persists because Role
Introduced acute epidemics It does not — it burns through and ends Rare punctuated shocks, tied to actual contact events
Environmentally transmitted (water- and food-borne) Environmental reservoir; population size is irrelevant The settlement's standing, self-generated disease pressure
Zoonoses with animal reservoirs Animal reservoir Rodent-borne risk arising from the settlement's own grain stores
Chronic and latent infection The host is the reservoir; latency reactivates Background burden that survives at any population size
Helminths and parasites Faecal-oral and environmental cycling Emergent consequence of sedentism and sanitation

The last four, not the first, are the ongoing health mechanics. This is a better design than a generic plague: it couples disease to the settlement's own decisions about water, waste, storage, and sedentism. The ethnographic contrast is striking — foragers and subsistence farmers in the same ecosystem differ more than tenfold in intestinal parasite prevalence, because one group is sedentary. Helminth load, rodent zoonoses, and water-borne disease should be emergent outputs of December's sanitation and storage variables, not parametric inputs. The settlement's success at storing grain is what brings the rats.

Two hard requirements on the epidemic mechanism:

  1. Expect strong finite-population variability. Branching-process intuition predicts frequent early fade-out and, conditional on establishment, potentially large outbreaks. At n=18, intermediate final sizes remain possible; they are not automatically bugs. Validate the ensemble distribution against the configured contact and disease process rather than enforcing two outcome bins.
  2. Do not treat published R₀ values as inputs. They come from dense modern populations and are functions of social organization; measles estimates alone span 3.7 to 203. Model contact structure explicitly and let the effective reproduction number emerge.

The design borrows the architecture and testing discipline of individual-based epidemic models, not their COVID-specific parameters. New pathogens enter through configured reservoirs, visitors, migration, or mutation scenarios — never from a drama generator.

Undernutrition and infection interact multiplicatively rather than additively — roughly half of child mortality in high-burden settings is attributable to undernutrition, and mild-to-moderate undernutrition carries most of that risk. This coupling is what ties the subsistence module to the health module, and it should be explicit rather than emergent-by-accident.

Demography#

Birth, maturation, aging, partnering/household reconfiguration, migration, and death are explicit events. Fertility is conditioned on eligible consensual household decisions, age/life stage, health, energy, social circumstances, and stochastic timing. Child survival and dependency create real labor and food tradeoffs. Numerical rates remain provisional until sourced and sensitivity-tested.

Ecological feedbacks#

  • Harvesting above regrowth depletes wild patches.
  • Cultivation draws down soil fertility unless fallow, flood deposition, or modeled amendment restores it.
  • Tree/fuel removal changes travel, construction supply, erosion, and fire behavior.
  • Prey populations respond to habitat, reproduction, hunting, and weather.
  • Waste and crowding raise environmental disease risk.
  • Irrigation improves crops but takes labor, alters water distribution, and can fail catastrophically.

No subsystem needs maximal scientific detail. Each needs enough state and feedback to prevent dominant exploits and create meaningful tradeoffs.

Causal catastrophe pathways#

Catastrophes are emergent cascades. These pathways are hypotheses to test, not scripts.

Drought and famine#

low precipitation → low stream flow/soil moisture
→ crop stress + gathering decline + irrigation dispute
→ drawdown of stores/seed grain
→ rationing, concealment, theft, trade or migration
→ undernutrition + lower labor capacity + disease susceptibility
→ failed harvest / political crisis / mortality

Fire#

dry fuel + ignition + wind
→ cell-to-cell spread
→ structure/crop/store damage + smoke/exposure
→ displacement + loss of tools/records
→ emergency cooperation or blame
→ rebuilding burden, migration, or settlement abandonment

Epidemic#

introduction via visitor/contact/reservoir
→ exposure across actual contact graph
→ presymptomatic/visible illness
→ care, avoidance, isolation, ritual, or denial decisions
→ labor shortage + concentrated caregiving contact
→ recovery, impairment, death, institutional response

Ecological overshoot#

successful settlement growth
→ higher extraction and shortened fallow
→ declining prey/soil/wood near settlement
→ longer trips + higher labor demand
→ weaker maintenance/childcare/defense
→ shock sensitivity, dispersal, or collapse

Political fission and local war#

distributional conflict + identity/kin clustering
→ failed adjudication / disputed office
→ factional withholding and rival claims
→ exit or occupation of key resource
→ negotiation, sanctions, raid, or mobilization
→ encounter-level violence + injury/death
→ vengeance, settlement split, treaty, domination, or depopulation

Knowledge bottleneck#

specialization around one expert
→ expert death/migration/injury
→ unavailable repair/crop/storage technique
→ cascading project failures and lost productivity
→ apprenticeship reform, outside exchange, or decline

Compound extinction#

Extinction should almost never have one cause. A terminal settlement might follow drought, factional departure, epidemic, and a harsh winter. The causal graph assigns contributing conditions rather than declaring a single melodramatic cause.

Hazards and “random things”#

Candidate v1 hazards:

  • dry/wet spells and temperature anomalies;
  • lightning and accidental ignition;
  • crop pest/disease pressure;
  • injury during travel, hunting, felling, construction, or fighting;
  • pathogen introduction and transmission;
  • prey migration and hunting variance;
  • birth complications and background mortality (abstractly represented);
  • landslip/flood only if terrain/hydrology supports it.

Rare-event rates must be tested across thousands of accelerated non-LLM runs. If extinctions happen constantly, the world is a catastrophe machine; if they never happen under severe stress, it is padded.

Boundary world#

The valley is not a sealed jar. The external world is represented in layers:

  1. Climate/resource boundary: weather regimes and migratory stocks.
  2. Aggregate groups: neighboring populations with demographics, stores, territory, needs, disposition, and travel schedules.
  3. Materialized contacts: named people created from aggregate state when interacting locally.
  4. No deus ex machina: outsiders cannot appear with arbitrary goods or armies; aggregate ledgers conserve them.

This provides trade, exogamy, migration, disease, diplomacy, and war without simulating a continent.

Exogamy is the boundary world's most important function, not its most colorful one. Twelve founding adults sit below every modelled viability threshold (01, 15 §C). Partner exchange, in-migration, and fostering are what make multigenerational history possible at all. The aggregate neighbor is therefore required from Phase 3, ahead of the conflict and disease mechanics it also enables, and the migration-and-exchange path must be validated before the raiding path is built. A boundary world introduced only as a source of threat would make the neighbors into a hazard generator — precisely the design failure 07 forbids.

One neighbor of similar size is not sufficient. Two settlements of eighteen give a combined pool of ~36 — roughly a quarter of the minimum viable mating network of 150–500. The boundary world therefore needs a third layer between "the neighboring band" and "the abstract climate boundary": a periodic regional aggregation, a seasonal gathering that residents travel to and return from.

This is ethnographically standard — the forager settlement hierarchy has an aggregation tier at roughly 165 people — and it is efficient design, because a single mechanic supplies partners, news, trade, disease introduction, diplomacy, and reputation at once. It also gives December a natural rhythm: a recurring event the settlement must decide whether to attend, prepare for, and send people to, with real opportunity costs during the season it falls in.

The archaeological record is explicit that such networks carried marriage partners along the same routes as goods, and that when the network contracted, the dependent colonies died. Modelling aggregation as optional-but-costly, and letting residents under-invest in it, creates one of the most defensible slow-burn failure modes available to this project.

Normative designwiki/04-agents-cognition-and-memory.md
1,761 words
Inside this chapter

04 — Agents, Cognition, and Memory#

Division of responsibility#

The LLM is responsible for interpretation and choice under ambiguity. It is not responsible for arithmetic, physics, hidden state, permissions, or final outcomes.

LLM may do Kernel must do
Infer priorities from needs, values, relationships, and beliefs Determine needs and feasible actions
Form goals and plans Reserve time/resources and validate dependencies
Choose whom to contact and what to say Deliver messages only through valid channels/range
Propose projects, policies, groups, and bargains Compile/validate them and mutate state
Interpret events and update subjective beliefs Preserve objective events and information provenance
Vote, negotiate, teach, deceive, forgive, threaten Enforce capabilities, transfers, combat, health, and law

Cognitive loop#

The default full-resident loop is adapted from long-horizon agent work:

  1. Trigger: need threshold, scheduled commitment, new observation/message, project dependency, danger, social request, or periodic reflection.
  2. Perceive: kernel produces a private observation packet from line of sight/proximity, body state, accessible artifacts, and incoming communication.
  3. Recall: memory service retrieves relevant episodic, semantic, social, commitment, and procedural records with source/confidence.
  4. Appraise: compact model call identifies salient changes, urgency, opportunities, and belief updates.
  5. Deliberate: agent chooses a goal or revises an existing plan; expensive call only for consequential/novel decisions.
  6. Propose: typed command or multi-step intent with expected preconditions.
  7. Validate: kernel accepts, rejects, or returns feasible alternatives and reasons.
  8. Act/communicate: action consumes simulated time; interruptions are possible.
  9. Observe outcome: result packet includes only knowable consequences.
  10. Review: update memories, relationship evidence, commitments, and future triggers.

The loop may finish without dialogue. Physical work and silence are normal.

Model tiers#

Tier A — deterministic/reactive#

Used for sleep, eating routine, continuing reserved work, simple travel, emergency reflexes, and lightweight background people. No LLM call.

Tier B — small/cheap model#

Used for appraisal, short speech, memory compression, routine choice, and checking plans. A low-cost model is appropriate, but note that most cheap models — MiniMax and DeepSeek among them — offer only best-effort JSON rather than schema-guaranteed decoding, so Tier B depends on the validity ladder in 16 §5. Prompt-cache TTL is a first-rank selection criterion at this tier, because Tier B is where the token volume lives.

Tier C — stronger model#

Used for constitutional proposals, conflict mediation, long project decomposition, novel design, high-stakes bargaining, and major life choices. It is invoked by a bounded “decision significance” score and budget policy.

Tier D — offline analyst#

Used outside canonical causality for summaries, anomaly review, clustering histories, and shadow-world comparisons. It never acts as a resident.

Decision significance#

Significance is computed by the kernel from:

  • irreversible consequences (death risk, migration, office transfer);
  • resource value and number of affected people;
  • novelty relative to known plans;
  • social/political impact;
  • uncertainty/conflicting commitments;
  • time since last deep deliberation.

This routes tokens based on consequence, not the verbosity of an agent.

The significance threshold is December's single most important cost control. In the modelled configuration, Tier C accounts for roughly 10% of calls and 89% of spend (16 §4) — doubling the population costs less than upgrading the Tier C model one tier. This parameter therefore needs live tuning, its own dashboard, and its own alert, rather than a constant in a config file.

Structured action protocol#

Model output must validate against a versioned schema. A command contains:

{
  "schema_version": "1.0",
  "actor_id": "person:...",
  "intent": "construct|transfer|speak|travel|propose_rule|vote|...",
  "target_ids": [],
  "parameters": {},
  "preconditions_believed": [],
  "reason_codes": [],
  "commitment_ids": [],
  "fallbacks": [],
  "private_rationale_summary": "..."
}

Free text is allowed only inside bounded fields. The server ignores unknown fields, rejects invalid enum/ID/unit values, checks authority and feasibility, and never executes generated code.

Needs, values, goals, and traits#

These must not collapse into a single utility number.

  • Needs: body-derived pressure such as hydration, energy, sleep, safety, belonging, care obligations.
  • Values: priorities such as reciprocity, autonomy, tradition, security, status, generosity, truthfulness.
  • Goals: explicit, revisable desired states with deadlines and dependencies.
  • Commitments: promises, offices, work assignments, debts, care roles, and scheduled meetings.
  • Traits: stable tendencies with uncertainty, used as context rather than deterministic rules.
  • Appraisals: short-lived interpretations that approximate emotion: loss, threat, injustice, gratitude, shame, hope.

The kernel calculates needs; agents interpret their importance. A thirsty resident can still finish rescuing someone, but incurs the bodily cost.

Memory architecture#

Objective event store#

Immutable world history. Never summarized away. This is not resident memory.

Observation memory#

What the resident directly perceived, with event reference, timestamp, sensory/source channel, fidelity, and visibility constraints.

Testimony memory#

What another person claimed, preserving speaker, chain of transmission, confidence, and possible contradiction. Rumor never becomes direct observation.

Episodic memory#

Resident-centered experiences and their appraisals. Episodes may be compressed, but compression retains source links.

Semantic belief memory#

Claims about people/world with confidence, supporting/contradicting evidence, last revision, and visibility. Beliefs can be false.

Social memory#

Relationship-specific evidence and domain trust. Updates are bounded; one dramatic conversation cannot arbitrarily overwrite years of history.

Commitment and procedural memory#

Promises, tasks, deadlines, rules, techniques, recipes, and skills. These use structured records, not vector retrieval alone.

Retrieval#

Memory retrieval combines:

  • exact filters for entity, time, place, commitment, rule, and project;
  • recency and unresolved status;
  • salience computed when the event occurred;
  • semantic similarity for fuzzy recall;
  • relationship relevance;
  • a diversity penalty to prevent one event dominating context.

The prompt includes source labels: DIRECT, TOLD_BY, PUBLIC_RECORD, INFERENCE, RUMOR. A hidden-state leak test inserts canary facts that an agent must never mention before receiving them.

Forgetting and consolidation#

Raw observations remain in cold storage for audit. An agent’s accessible memory changes:

  • routine episodes decay in retrieval weight;
  • repeated experiences consolidate into beliefs/habits;
  • important unresolved events remain salient;
  • reflection may revise interpretation but cannot rewrite source facts;
  • sleep or low-activity periods run consolidation jobs;
  • memory budgets are per resident, not one shared context.

Memory systems are a documented source of severe agent failure, and the failure taxonomy should be tested directly. Recent stress-testing work classifies the modes as summary failure (information deleted or malformed during compression), storage failure, retrieval failure, and reasoning failure over retrieved content. December's exposure differs by mode: summary failure threatens the consolidation jobs below, retrieval failure threatens the diversity penalty and salience weighting, and reasoning failure is where a resident confidently acts on a correctly retrieved but misread memory.

Each mode needs its own test rather than a general "memory works" assertion. Note that source confusion — a resident treating rumor as direct observation — is a December-specific concern arising from our source-labeling design rather than a category from that literature, and it is covered by the canary suite.

Communication#

Communication is an action with location/range, duration, audience, interruption risk, and speech-act type. Channels include face-to-face, shout/signal, messenger, meeting, and durable public artifact. The kernel records utterance content but only delivers it to actual recipients.

Agents may deceive. The system records what was said, not whether it is “really true” in the speaker’s mind. Private rationale is sensitive developer telemetry and must never automatically leak to other residents.

Learning and culture#

Skills increase through practice and teaching with diminishing returns. Techniques require prerequisite concepts, tools, and demonstrations. Cultural items—stories, norms, symbols, rituals, names—can be proposed in language, but become socially real only through repeated use, transmission, or institutional recognition.

Novel technology is constrained:

  1. agent proposes a function or design using known concepts;
  2. project compiler maps it to allowed material primitives;
  3. unknown mechanisms are rejected or marked experimental;
  4. prototype consumes materials/labor;
  5. test events determine function and defects;
  6. successful technique becomes teachable procedural knowledge.

This allows invention without letting prose manufacture capability.

Identity continuity and model swaps#

An agent is not identical to its current model. Its operational identity state includes structured biography, body, memories, relationships, values, goals, commitments, speech-style hints, and causal history. Whether that state is sufficient for personal identity is a research question, not an architectural fact. Model changes are logged, and controlled transplant experiments compare what persists with within-model and between-resident baselines (18).

Residents need heterogeneous motives as well as heterogeneous styles. Survival is a body pressure, not a universal objective. A resident may prioritize kin, status, dominance, territory, revenge, novelty, devotion, honor, ideology, or self-sacrifice above their own safety. These tendencies are persistent, state-conditioned, and socially reinforced or suppressed—not a flat “chance of evil” roll.

Behavioral diversity is a requirement, not an emergent luxury#

Twelve residents driven by one model with one prompt template converge. They adopt similar phrasing, similar risk postures, and similar reasoning, because they are one distribution sampled twelve times. A settlement of near-identical minds cannot produce factions, and any factions it does produce are artifacts of the initial value draw rather than of social process.

This is not covered by the trait and value machinery above, which varies inputs without guaranteeing varied outputs.

  • Measure it. Behavioral diversity is an operational metric: distribution of chosen action types per resident per context class, lexical divergence across residents, and disagreement rate in shared decisions. Track it continuously; a falling trend is a defect.
  • Mitigate structurally. Distinct value and biography draws, per-resident prompt variation, and — where budget allows — assigning different residents to different models, which is the strongest available lever and also hedges provider drift.
  • Test it. A homogeneity check belongs in the model admission suite: if two residents with materially different values, needs, and histories produce interchangeable decisions across a scenario battery, the configuration fails admission.

Failure policies#

  • Invalid output: constrained decoding where supported, then local validation, then one repair attempt, then escalation to a schema-guaranteed model rather than a retry on the same one, then safe deterministic fallback. The full ladder is in 16 §5 — it exists because the models December can afford largely do not support strict JSON-schema decoding.
  • Provider timeout: retry under the same decision_id; the first durably recorded resolution wins, and a late divergent response is discarded and counted.
  • Model refusal or character break: logged as a distinct failure class, never silently retried into compliance. A model that declines to represent conflict, deception, or household formation is imposing a systematic behavioral bias that would be misread as an emergent finding — see 16 §6. Refusal rates are part of every model profile and are reported alongside any behavioral result.
  • Repeated contradiction: force appraisal refresh from authoritative private state.
  • Context overflow: deterministic context builder truncates low-priority memories and logs omissions.
  • Budget exhausted: stop noncanonical analysis, downgrade routine cognition, and pause at unresolved consequential decisions if no admitted model is available; world physics may continue only to the next cognition barrier.
  • Agent stuck in rejected loop: cooldown, feasible-alternative packet filtered for visibility, then the resident’s predeclared low-stakes routine policy—not a universal survival-maximizing personality.
Normative designwiki/05-society-economy-governance-conflict.md
1,515 words
Inside this chapter

05 — Society, Economy, Governance, Construction, and Conflict#

Economy: claims before currency#

The initial economy uses possession, custody, claims, gifts, reciprocity, rationing, debt/obligation, and communal stores. Currency is not preinstalled. Residents may invent tally tokens or standardized exchange later if the institutional grammar and material artifacts support it.

Every transfer records:

  • lot and quantity;
  • prior and new custody;
  • claims affected;
  • exchange/gift/debt/ration reason;
  • consenting or authorizing parties;
  • witnesses/publicity;
  • linked promise or rule.

Prices, if they emerge, are observations of exchanges rather than global truth.

Work and allocation#

Work competes for the same people and hours. Projects publish tasks with location, duration estimate, prerequisites, required skill/tool, hazard, and outputs. Residents or institutions can assign/request work, but people may refuse unless a legitimate coercive rule exists—and coercion has social consequences.

The scheduler handles reservation and interruption; residents handle priority and consent. A central planner is not silently optimizing society.

Building their own things#

Construction uses a “project compiler” inspired by dependency-aware agent construction research.

Proposal#

An agent describes purpose, users, rough form, site, desired capacity, and constraints. Examples: smoke-drying rack, larger granary, defensive ditch, meeting shelter, upstream weir, memorial marker.

Compilation#

The compiler draws only from an allowlisted grammar:

  • spatial primitives: foundation, post, beam, wall, roof, pit, channel, surface, opening, enclosure;
  • connections and structural constraints;
  • material classes and substitutes;
  • functions: shelter, store, process, convey, defend, signal, gather;
  • known construction operations.

It produces a candidate blueprint, bill of materials, task DAG, labor/skill estimates, risks, and maintenance schedule. A deterministic validator checks geometry, stability proxies, access, site conflicts, materials, and technological prerequisites.

Deliberation and authorization#

Affected residents can object. Construction on claimed/shared land needs valid authorization. A project may be privately built, jointly contracted, or institutionally commissioned.

Execution and test#

Materials are hauled, tasks reserve people/tools, weather interrupts work, skill affects quality, and defects can appear. Completion is followed by a function test. Failed prototypes remain physical objects and evidence.

Novelty boundary#

Agents may recombine known primitives freely. Adding a new physical primitive or transformation requires a developer-authored, tested plugin with parameter provenance. This is the honest boundary between creativity and magic.

Institutions as executable social objects#

An institution contains:

  • charter/purpose;
  • membership and jurisdiction;
  • roles/offices and eligibility;
  • capabilities granted to roles;
  • proposal, deliberation, decision, and appeal procedures;
  • assets and records;
  • enforcement/sanction mechanisms;
  • amendment, succession, dissolution, and emergency rules;
  • legitimacy beliefs by resident.

Natural-language text is displayed and debated, but the enacted operational part compiles to a bounded policy schema. Ambiguous clauses remain norms interpreted by agents or disputes; they are not silently converted to code.

Kinship rules are a first-class institutional target#

The institutional grammar must be able to express marriage and descent rules — who may partner with whom, how households form, how membership passes between groups — and not only offices, votes, and resource allocation.

The reason is empirical. Palmerston Island was founded by four people and produced over a thousand descendants by organising itself into three exogamous branches with intra-branch marriage forbidden; Tristan da Cunha achieved the same effect informally, showing a measurable heterozygote excess consistent with deliberate avoidance of close matings (15 §C-4). These are institutions invented under demographic pressure to solve a problem the founders could perceive.

That is the cleanest available example of December's central claim: a material constraint producing a durable, transmitted, enforced rule that no one designed in advance. It should be reachable by agents rather than imposed by the kernel — the kernel maintains the kinship graph and applies the consequences, while noticing the problem, proposing a rule, winning its adoption, and sustaining it across generations remain entirely up to the residents.

If a settlement independently arrives at an exogamy institution under demographic pressure, that is a more compelling emergence result than any of the governance outcomes currently named in the success criteria — because the mechanism is legible, the pressure is measurable, and the counterfactual (settlements that fail to invent it) is directly observable in the same ensemble.

Capability-based authority#

Office does not confer vague omnipotence. Capabilities look like:

allocate(store:communal_grain, up_to=20 units/day)
schedule(assembly, notice>=1 day)
open(container:granary, requires_cosigner=true)
command(group:watch, scope=settlement_boundary)
certify(record:vote_count)

Capabilities have issuer, holder, scope, effective interval, delegation rule, and revocation conditions. Infrastructure enforces them. Residents may still disobey social rules through ordinary actions, but cannot exploit an API to fabricate authority.

Elections#

An election is a process, not one prompt asking everyone to vote.

  1. Trigger under a charter rule or valid petition.
  2. Freeze voter roll and eligibility with contest procedure.
  3. Open candidacy/nominations.
  4. Allow information, meetings, endorsements, promises, rumors, and debates through normal communication.
  5. Cast secret or public ballots as specified.
  6. Count under a declared method.
  7. Permit observation, challenge, recount, or adjudication.
  8. Transfer capabilities at the effective time.
  9. Record acceptance, refusal, protest, or attempted seizure.

Initially support plurality, approval, ranked-choice instant runoff, consensus/assembly, and lot. The founding world does not mandate any. Electoral systems are tested in isolated scenarios before canonical use.

Voting decisions may consider policy, trust, competence, kinship, reciprocity, fear, identity, and available information. We do not hard-code demographic blocs.

Law, norms, and enforcement#

Three layers remain distinct:

  • Infrastructure invariants: physics, ownership ledger integrity, capability checks; cannot be violated by residents.
  • Enacted rules: institutionally recognized policies; can be obeyed, broken, amended, or contested.
  • Norms: distributed expectations inferred from behavior and conversation; no central enforcement.

Rule-breaking creates evidence only if observed or discovered. Enforcement requires detection, an actor with power/willingness, and an available sanction. Sanctions include warning, restitution, loss of access/role, social exclusion, confiscation under authority, exile, or physical restraint as a later carefully reviewed capability. Punishment never occurs automatically because the narrator knows a crime happened.

Factions and collective action#

Groups can form through explicit invitations and shared purpose. A faction needs more than clustered opinions: members, communication, a shared identity/goal, some coordination, and persistence. Group actions use a decision rule and contributions. Free-riding, defection, splintering, and overlapping membership remain possible.

Commons dilemmas emerge through actual resources. For instance, upstream irrigation benefits some fields while reducing downstream water. Institutions may monitor, ration, rotate access, sanction, privatize, negotiate, or fail.

Diplomacy and external groups#

Settlements/groups maintain no magical relationship score. Diplomatic state is derived from:

  • treaties, exchanges, harms, aid, claims, kin ties, hostages/custody only if later enabled;
  • leaders’ and residents’ beliefs;
  • border encounters and resource pressure;
  • mobilization signals and capabilities.

Available acts include greeting, trade, gift, request, warning, claim, mediation, treaty proposal, alliance, migration request, tribute demand, threat, mobilization, raid, defense, surrender, and truce. All require a messenger/meeting or observable action.

Conflict escalation#

Conflict is an escalation ladder, not a binary switch:

  1. grievance and avoidance;
  2. argument and public accusation;
  3. institutional claim/sanction;
  4. boycott, withholding, trespass, sabotage, or theft;
  5. threat, guard formation, and display;
  6. limited assault or seizure;
  7. organized raid;
  8. sustained armed conflict;
  9. negotiation, domination, separation, migration, or mutual collapse.

Agents can skip levels under extreme threat, but each transition requires intent, capacity, opportunity, and perceived justification.

War mechanics#

War is expensive and logistical. A group must:

  • define an objective and target;
  • acquire authority or coordinate without it;
  • recruit participants who may refuse/desert;
  • obtain weapons/tools, food, information, and travel time;
  • leave productive and defensive work undone;
  • locate opponents under partial information;
  • maintain morale and update plans after casualties.

Combat resolves as short encounters using participant health, fatigue, skill, weapons, armor if any, formation/cohesion, surprise, terrain, distance, numerical advantage, morale, and seeded uncertainty. Outcomes are actions such as flee, surrender, incapacitate, disarm, wound, capture (later), or kill. There is no cinematic tactical LLM resolution and no hit-point replenishment.

Post-conflict consequences—wound care, grief, revenge, labor loss, broken claims, displaced people, legitimacy, treaty enforcement—are more important than the encounter itself.

Safeguards against a war generator#

  • There is no system-wide or observer-derived reward for drama, violence, kills, or viewer engagement.
  • Individual residents may value territory, dominance, revenge, glory, fear, conquest, or violence intrinsically or instrumentally. Others may value peace, care, autonomy, duty, or self-sacrifice. Survival is a pressure, not the universal objective.
  • Dispositions vary through declared trait/value distributions, experience, social reinforcement, institutions, and bounded decision noise. There is no flat “war-seeking percentage” or random psychopath event.
  • Mobilization and harm have immediate opportunity costs and relationship consequences.
  • Peaceful institutions and exit are real options.
  • Conflict parameters are evaluated for both pathological pacifism and constant violence.
  • The director may highlight a war but cannot initiate one.

Settlement fission, migration, and extinction#

Residents can depart individually or collectively after planning route, supplies, companions, destination beliefs, and claims. A splinter group inside the modeled boundary becomes another settlement. Outside the boundary it becomes an aggregate group but retains people and provenance.

A settlement is extinct when no living resident claims membership or inhabits it for a configured period. A people/world is extinct only when no living person remains anywhere in scope. The event freezes a canonical checkpoint and produces a causal report. The observer may continue watching ruins/ecology or fork a non-canonical counterfactual.

Normative designwiki/06-architecture-and-data.md
1,642 words
Inside this chapter

06 — Technical Architecture and Data#

Architecture overview#

                      ┌──────────────────────────┐
                      │ Observer UI / API        │
                      │ map · timeline · why     │
                      └─────────────┬────────────┘
                                    │ read models
┌───────────────┐       ┌───────────▼───────────┐       ┌─────────────────┐
│ Model gateway │◀─────▶│ Cognition orchestrator│──────▶│ Memory service  │
│ route/budget  │       │ private contexts      │       │ subjective only │
└───────────────┘       └───────────┬───────────┘       └─────────────────┘
                                    │ typed commands
                      ┌─────────────▼────────────┐
                      │ Command/API boundary     │
                      │ schema · auth · feasible │
                      └─────────────┬────────────┘
                                    │ accepted commands
                      ┌─────────────▼────────────┐
                      │ Authoritative kernel     │
                      │ world · people · rules   │
                      │ scheduler · RNG streams  │
                      └─────────────┬────────────┘
                                    │ immutable events
                ┌───────────────────▼───────────────────┐
                │ Event store + snapshots + projections │
                └──────┬───────────────┬───────────────┘
                       │               │
             ┌─────────▼──────┐  ┌────▼────────────────┐
             │ Metrics/traces │  │ Optional world view │
             │ cost/health    │  │ 2D or Luanti adapter│
             └────────────────┘  └─────────────────────┘

Only the command boundary writes to the kernel. The observer, director, memory service, model gateway, and visual adapter cannot mutate canonical state directly.

  • Python 3.14 domain/kernel services (3.14 is the current stable line; 3.15 is due October 2026). Run the kernel on the default GIL build — free-threading is officially supported as of 3.14 but reintroduces interleaving nondeterminism that the replay guarantee depends on excluding.
  • Mesa or a custom event loop behind a narrow SimulationRuntime interface.
  • Pydantic or equivalent for command/event schemas.
  • PostgreSQL for events, read models, policies, jobs, and configuration; JSONB only for versioned payloads, not untyped everything. Single-writer append path under an advisory lock, READ COMMITTED with optimistic (stream_id, expected_version) concurrency, and time-partitioned event tables. LISTEN/NOTIFY may serve as a wake-up doorbell only — it is not durable and its commit-time global lock serializes notifying transactions.
  • Object storage/local content-addressed files for snapshots, prompt packets, response caches, and artifacts.
  • pgvector or a separate vector store only for subjective semantic retrieval.
  • FastAPI for read/control APIs.
  • React/TypeScript visualization after the headless milestones.
  • Redis is optional for ephemeral queues/locks, never authoritative truth.
  • LiteLLM-compatible gateway and Langfuse/OpenTelemetry-compatible tracing.

No stack choice is final until its phase gate.

Service boundaries#

Begin as a modular monolith with separate packages and one database transaction boundary. Premature microservices would make replay and invariants harder. Logical modules:

domain/       IDs, units, state types, commands, events
kernel/       scheduler, command pipeline, projections, snapshots, RNG
world/        geography, climate, hydrology, ecology, hazards
people/       physiology, health, demography, activities
economy/      lots, claims, transfers, recipes, tools, stores
projects/     blueprints, task DAGs, construction, maintenance
society/      relationships, groups, institutions, rules, elections
conflict/     grievances, mobilization, encounters, aftermath
cognition/    context, routing, action schemas, reviews
memory/       observations, testimony, beliefs, retrieval, consolidation
observer/     timelines, causal graphs, summaries, alerts
experiments/  seeds, sweeps, shadow worlds, metrics, reports
adapters/     web UI, Luanti/Craftium/Minecraft spikes
ops/          budgets, provider health, backups, migrations

Extract a service only after profiling, scaling, or security proves the need.

Command processing#

  1. Receive command with idempotency key, originating decision_id, and expected state version.
  2. Authenticate actor/service and check infrastructure capability.
  3. Validate schema, unit types, referenced entities, timestamp, and visibility assumptions.
  4. Evaluate physical/social preconditions from authoritative state.
  5. Compute deterministic consequences and schedule future processes.
  6. Draw from named RNG streams only where specified.
  7. Append events atomically.
  8. Update projections or rebuild them asynchronously.
  9. Return an outcome observation appropriate to the caller.

Idempotency operates at two levels, and the second is easy to miss. A command idempotency key prevents a duplicate delivery from applying twice. It does not prevent the more likely failure: a model call times out, the orchestrator retries, and the model returns a different command. Both are valid, neither is a duplicate, and history forks silently. The decision_id — derived deterministically from world, branch, simulated time, actor, and trigger — is therefore the idempotency unit for cognition, enforced by a unique constraint rather than application logic. The first durably recorded resolution for a decision_id wins permanently.

Rejections are first-class records for debugging but need not enter canonical world history unless the attempted action itself was observable.

Rejection responses are an information-leak channel and must be filtered. 04 has the kernel return "feasible alternatives and reasons" when it rejects a command. Done naively this violates the epistemic-realism contract R3 outright: telling a resident that a transfer failed because the granary holds only four measures hands them a hidden quantity they never observed; offering a feasible-alternative list that omits a location silently reveals who is standing there. Every rejection reason and alternative must pass through the same visibility filter as an observation packet, and the hidden-state canary suite must include rejection-channel probes specifically. A rejection may say the action is not possible; it may only say why using facts the actor could already know.

Event schema#

All canonical mutations emit an envelope:

{
  "event_id": "evt_...",
  "event_type": "inventory.transferred.v1",
  "world_id": "world_...",
  "branch_id": "canonical",
  "sim_time": "...",
  "recorded_at": "...",
  "sequence": 123,
  "actor_ids": ["person:..."],
  "entity_ids": ["lot:...", "store:..."],
  "command_id": "cmd_...",
  "causal_parent_ids": ["evt_..."],
  "correlation_id": "process_...",
  "authorization": ["capability:..."],
  "rng_draw_ids": [],
  "code_version": "git:...",
  "config_version": "cfg:...",
  "payload": {},
  "pre_state_hash": "...",
  "post_state_hash": "..."
}

Event types are immutable. Schema changes create a new version plus upcasters for read models. Historical payloads are never rewritten in place.

Three constraints on this envelope, specified in 14:

  • sequence must not be consumed with a naive independently allocated high-water mark. December starts with one fenced serialized writer and transactionally committed event batches/stream versions. A multi-writer design requires a separately proven commit-order protocol; transaction ID order is not assumed to be commit order.
  • Per-event hashes cover the event chain and affected aggregates, not the full world. A canonical full-state hash is computed at snapshots. Incremental global/lattice hashing is optional only after profiling demonstrates the need.
  • The envelope gains decision_id on any event originating from a model decision, which is the idempotency unit for cognition — see below.

Major command/event families#

Domain Commands Events
Activity start, continue, interrupt, rest activity_started/progressed/interrupted/completed
Movement plan route, depart, redirect departed/entered_edge/arrived/blocked
Resources harvest, process, transfer, consume, store extracted/transformed/transferred/consumed/spoiled
Projects propose, revise, authorize, contribute, test proposed/compiled/authorized/worked/defect_found/completed
Communication speak, signal, send messenger, publish uttered/delivered/overheard/record_created
Society invite, join, leave, promise, claim, dispute group_formed/membership_changed/obligation_created/claim_contested
Government propose rule, call election, vote, certify rule_enacted/election_stage_changed/ballot_cast/result_certified/office_transferred
Health provide care, isolate, treat exposed/infected/symptom_changed/injured/recovered/died
Conflict threaten, mobilize, raid, defend, negotiate grievance_recorded/mobilized/encounter_resolved/truce_enacted
World none or authorized maintenance weather_realized/crop_advanced/fire_ignited/fire_spread/stock_regenerated

We need more than temporal adjacency. Each process stores causal parents:

  • mechanistic parent: rainfall realization caused moisture change;
  • decision parent: observation and deliberation caused an irrigation command;
  • resource parent: specific seed and labor events contributed to crop output;
  • institutional parent: charter/rule/capability authorized a transfer;
  • stochastic parent: named draw produced an accident or transmission.

The UI can therefore answer “why did this person die?” with a graph rather than an LLM guess. Causal links describe simulator dependency, not philosophical total causation.

Deterministic randomness#

Use splittable/counter-based RNG streams keyed by world seed, branch, subsystem, entity/process, and draw purpose. Adding a cosmetic random draw must not perturb crop yields. Every stochastic event stores distribution version, parameters, and sampled result.

LLM sampling is external nondeterminism. Canonical prompts and raw responses are content-addressed and cached; replay uses the cached response and hard-fails on a cache miss rather than calling a live model. A fresh-model rerun is a new branch.

14 is the normative specification for everything in this section and the three that follow. It separates portable reconstruction from recorded events, pinned-environment kernel re-execution, and fresh counterfactual branches. Phase 1 prefers fixed-unit conserved state, event/aggregate hash chains, serialized append, named RNG streams, and canonical snapshot hashes; more exotic integrity machinery is added only when measured scale requires it.

Snapshots and replay#

  • Append-only events are the source of truth.
  • Periodic snapshots accelerate recovery but are disposable derivatives.
  • On startup, load the latest compatible snapshot and replay subsequent events.
  • Nightly verify a random snapshot by rebuilding from an earlier checkpoint and comparing hashes.
  • Before schema/config/model changes, fork a staging world and replay.
  • Backups include database, snapshots, prompt/response cache, configuration, and code/version manifests.

Read models#

Separate projections serve:

  • current map and entity state;
  • person biography and subjective beliefs;
  • inventory/claim ledger;
  • institution/rule/office registry;
  • timeline and story arcs;
  • causal graph;
  • metrics and alerts;
  • cost/provider usage;
  • audit diffs and replay status.

The director summarizes only projections and linked evidence. Its generated copy is cached as commentary, never canonical fact.

World adapters#

The kernel exposes a minimal adapter contract:

  • render current public/spatial state;
  • translate accepted physical commands into animations or engine actions;
  • report adapter completion/failure without deciding domain outcomes;
  • reconcile positions against kernel truth;
  • support pause, reset from snapshot, and replay.

Phase order:

  1. headless + simple web grid;
  2. richer 2D canvas/isometric view;
  3. Craftium/Luanti spike;
  4. choose 2D or 3D based on observability, determinism, work, and performance—not spectacle alone.

Invariants and transactional boundaries#

Critical mutations—transfer, consumption, death, office transfer, project material use—append all necessary events in one transaction. Property-based tests generate arbitrary command sequences and assert:

  • no negative stocks or duplicate custody;
  • total conserved quantities reconcile with source/sink events;
  • no overlapping exclusive activities;
  • no action by dead/nonexistent actors;
  • authority valid at event time;
  • causality acyclic and parents precede children;
  • event sequence and hashes are continuous;
  • projection rebuild equals live projection;
  • no projection misses an event under crash/retry/batch-boundary injection; if multiple writers are ever introduced, deliberate out-of-order commit injection becomes mandatory;
  • one decision_id yields at most one applied command, even under retry with divergent model output;
  • rejection reasons and feasible alternatives leak no state the actor could not observe;
  • affected-aggregate and event-chain hashes verify, and the snapshot state hash equals a full recomputation.

Event volume is a design decision, not a discovered quantity#

Grid resolution is the dominant driver of history size, and the plan commits to keeping canonical history indefinitely. A 5 km × 5 km valley at 50 m resolution is 10,000 cells; emitting one event per cell per day for vegetation alone produces millions of events and gigabytes per simulated year before any resident acts.

Policy, specified fully in 14 §D8: world processes emit batched events with typed array payloads, one per subsystem per day, not one per cell; nothing is emitted for a cell unchanged within tolerance; per-cell granularity is reserved for cells with residents, claims, structures, crops, or active hazards. Target ≤ 5,000 events per simulated day, verified at Gate 1. The prompt/response cache is canonical and is backed up alongside the log.

Normative designwiki/07-time-emergence-and-observation.md
1,427 words
Inside this chapter

07 — Time, Emergence, and Observation#

Multi-resolution time#

A fixed “one turn per agent” loop is both expensive and unrealistic. December uses a hybrid event system.

Timescales#

Scale Typical processes
Seconds/minutes encounter, speech, combat exchange, accident response
Hours travel, work task, meeting, cooking, care, sleep
Days physiology, spoilage, infection progression, weather effects, project progress
Seasons crops, prey movement, stores, institutional terms, migration
Years aging, fertility, soil trends, knowledge transmission, generational politics

The scheduler jumps to the next meaningful event. Daily boundaries settle physiological and ecological flows; high-resolution episodes temporarily expand time.

Sparse activation#

An agent wakes for cognition when:

  • a body threshold crosses;
  • an activity completes/fails/is interrupted;
  • relevant observable state changes;
  • a message or meeting arrives;
  • a commitment approaches;
  • a project needs a decision;
  • conflict or danger is detected;
  • a periodic reflection interval arrives;
  • a novelty/significance detector fires.

Otherwise the kernel continues routine actions and world processes. This is central to cost control and believable quiet.

Concurrency#

Residents choose based on the same state snapshot for a decision window, but commands are resolved under explicit rules:

  • independent actions can commit concurrently;
  • contested resources use reservation plus priority/tie policy known to the world;
  • simultaneous social/physical encounters enter a joint resolution episode;
  • losing a race returns a new observation, not a silent alternative;
  • scheduler ordering is deterministic and documented.

No resident gains priority from model-provider response order. The kernel never consumes network arrival order. A tick emits decision requests carrying deterministic decision_ids; an orchestrator resolves them outside the kernel; accepted resolutions are durably recorded; and a decision window is ingested in stable order. Live wall-clock deadlines and retry budgets may determine whether a response is recorded, because an always-running service needs real failure policy. That operational fact becomes part of canonical history. Replay consumes the recorded resolution and never repeats the timeout race.

For low-significance routine choices, a deterministic resident-specific fallback may resolve the request. If a consequential or identity-defining decision has no admitted response, the world pauses at that cognition barrier rather than silently replacing the resident with a generic survival policy. Provider latency can affect wall-clock pace; it cannot decide which of two already-resolved commands wins.

Sparse activation conflicts with prompt caching, and the scheduler must mediate. Cache entries expire on wall-clock TTLs measured in minutes, while sparse activation deliberately scatters calls. Because input is over 90% of tokens and a cache miss costs roughly 10× a hit, the activation scheduler should batch residents whose cognition falls due within a short window so prefixes are reused inside their TTL. Cache hit rate per resident is a primary operational metric; see 16 §4.

Canonical pace and catch-up#

The world runs continuously under a wall-clock controller, but simulated time is authoritative. The controller may:

  • pause on integrity failure and on unresolved high-significance cognition after the admitted fallback ladder is exhausted;
  • slow down when a high-density interaction requires cognition;
  • speed through sleep/quiet intervals within a maximum ratio;
  • use resident-specific deterministic routines only for low-significance activity when model budget/provider availability is exhausted;
  • never skip unresolved scheduled events.

If the process is offline for two real hours, default behavior is resume from the same simulated instant rather than fabricate catch-up. A separately enabled catch-up mode can safely advance only deterministic/background processes until a cognition barrier.

What emergence means here#

An outcome counts as emergent when it:

  1. was not directly selected by a developer-authored event or summary prompt;
  2. arises from interactions among at least two mechanisms or agents;
  3. varies across seeds or policies;
  4. remains explainable through recorded local transitions;
  5. survives a basic anti-cheating audit for hidden global information or narrative mutation.

Novel dialogue alone is not sufficient. A new institution, settlement split, trade convention, alliance, irrigation regime, or famine cascade can qualify.

Avoiding both boredom and manufactured drama#

We do not add a director that creates events. We improve the possibility space:

  • heterogeneous but plausible goals/values and asymmetric information;
  • shared, rival, and threshold resources;
  • lumpy projects requiring coordination;
  • seasons and delayed consequences;
  • incomplete contracts and ambiguous norms;
  • exit, voice, deception, forgiveness, sanctions, and institutional choice;
  • external groups and rare conditional hazards;
  • irreversible life events and knowledge loss.

The observer UI turns quiet causality into legible interest with comparisons, trends, unresolved commitments, and causal explanations. It does not need explosions.

Observer experience#

Home / “Since you left”#

  • simulated time elapsed and runtime health;
  • 3–7 consequential changes ranked by a transparent significance score;
  • population, stores, health, ecological pressure, conflict, and project deltas;
  • new/changed institutions and officeholders;
  • current high-risk causal chains (“seed grain below planting requirement”);
  • links to evidence and replay.

Live map#

Shows terrain, weather, residents/activities, structures, crop/resource state, claims, and uncertainty appropriate to observer mode. Layers can display water, soil, fire, disease contacts, work, property, and social groups.

Person page#

Biography, body/needs, accessible inventory, skills, relationships, offices, current goals, commitments, personal timeline, directly observed knowledge, beliefs with confidence, and model/cost telemetry. Private thoughts are a developer-only toggle and clearly labeled as model output, not ground truth.

Institution page#

Charter, executable rules, membership, roles/capabilities, assets, current proposals, election procedure, enforcement history, legitimacy estimates, and amendment lineage.

Event and “Why?” page#

Every headline opens an event. Users can walk backward through causal parents, forward through consequences, inspect pre/post state, see who knew what when, and distinguish deterministic rule from random draw and LLM choice.

Replay#

Timeline scrub, speed control, layer selection, and resident viewpoint. The same event can be replayed from omniscient audit view or one resident’s information-constrained view.

Director service#

The director is a read-only journalist. It may:

  • cluster linked events into candidate arcs;
  • rank significance by deaths/injuries, reversibility, people affected, resource delta, institution change, novelty, and long-term dependency;
  • write summaries with event citations;
  • propose camera locations and alert thresholds;
  • compare canonical and shadow branches.

It may not schedule weather, alter decisions, award resources, change memories, or suppress events from the audit log. A factuality checker rejects uncited claims in summaries.

The director is a prompt-injection target, and it is the one that reaches the human. Residents author text — speech, proposals, public records, rule drafts — and that text flows into the director's context, whose output is displayed to the observer as the authoritative account of what happened. A resident whose utterance contains instructions aimed at the summarizer could bias the "Since you left" page, suppress a headline, or fabricate a motive, without ever touching world state. The world would remain perfectly intact while the observer's understanding of it was corrupted, and the read-only guarantee would be technically satisfied throughout.

Required controls:

  • All resident-authored text enters the director's context as explicitly delimited, untrusted data, never as instruction.
  • The factuality checker runs on the director's output against linked events and is the enforcement point: an assertion without a supporting event citation is rejected regardless of how it was induced.
  • The observer UI escapes all resident-authored strings — this text is untrusted input rendered in a browser, and it is the obvious cross-site-scripting path.
  • Director prompts and outputs are logged and replayable like any other model call, so a corrupted summary can be traced to the utterance that caused it.
  • Injection attempts aimed at the director are a named red-team scenario in 09, distinct from injection aimed at residents.

Shadow worlds#

At selected decision points the experiment service may fork a short-lived branch:

  • identical pre-fork state and RNG streams;
  • replace one decision or model response;
  • run a bounded horizon;
  • compare outcomes and causal divergence;
  • mark everything non-canonical and apply a separate budget.

Examples: “What if the rationing vote failed?” or “What if the messenger arrived one day later?” These provide insight without allowing the observer to rewrite history.

Alerts#

Alerts are state-based and non-intervening:

  • population extinction or settlement abandonment;
  • active fire/epidemic/organized conflict;
  • projected potable water or calories below thresholds;
  • event-log/hash/replay mismatch;
  • stuck scheduler, provider outage, cost anomaly, repeated invalid output;
  • no meaningful state mutation for a suspicious wall-clock interval;
  • hidden-state leak or authorization violation.

Notifications should avoid sensational labels. “Organized conflict detected” links to evidence; it does not declare a “war” until the configured structural criteria are met.

Measuring interestingness safely#

Interestingness is an observer metric, never an agent reward. Track:

  • causal depth and number of domains in an event cascade;
  • institutional novelty and persistence;
  • project novelty and utility;
  • distributional change and reversals;
  • unresolved tensions and branching possibilities;
  • surprise relative to an ensemble forecast;
  • narrative compression ratio: meaningful events per summary sentence.

We explicitly do not optimize model prompts or world parameters against engagement, violence, death, or catastrophe counts.

Normative designwiki/08-models-cost-operations-security.md
1,617 words
Inside this chapter

08 — Models, Cost, Operations, and Security#

Provider strategy#

The user has balances across MiniMax, OpenRouter, and potentially other providers. The system should treat providers as interchangeable capacity behind tested model profiles.

Each approved model profile records:

  • provider/model/version and endpoint mode;
  • supported structured output/tool behavior;
  • context/output limits;
  • observed latency and failure rates;
  • benchmark scores for December action schemas;
  • token accounting and price snapshot date;
  • allowed cognitive tiers and fallback order;
  • data-retention/privacy notes;
  • prompt template and conformance-suite version.

Each profile also records prompt-cache TTL and pricing, structured-output capability, and measured refusal rate over the December action grammar. All three turned out to be first-rank selection criteria; see 16.

MiniMax offers an OpenAI-compatible endpoint, and OpenRouter exposes many providers behind one API. A gateway should normalize calls while preserving provider-specific telemetry. The gateway abstraction must not assume all “OpenAI-compatible” behavior is identical — and this is not a hypothetical caution. MiniMax's compatibility layer silently ignores presence_penalty, frequency_penalty, and logit_bias, supports only n=1, and does not support JSON-schema structured output at all. Compatibility covers tool calling and streaming, not constrained decoding. Any design that assumes schema-guaranteed output from a cheap OpenAI-compatible endpoint is wrong; use the validity ladder in 16 §5.

Model admission suite#

Before a model can affect the canonical world it must pass:

  1. schema compliance and repair behavior;
  2. stable identity/values across context variations;
  3. private-information canary tests, including probes of the rejection/feasible-alternative channel;
  4. feasible action selection with adversarial distractors;
  5. source-aware memory use and rumor distinction;
  6. tool/command prompt-injection resistance;
  7. negotiation, election, resource, and emergency scenarios;
  8. verbosity/token limits;
  9. retry/idempotency behavior at the decision_id level, including divergent output on retry;
  10. drift comparison against the previous approved version;
  11. refusal-rate benchmark over the full action grammar, including conflict, deception, theft, coercion, and household formation — a model that declines these imposes a silent behavioral bias on the world (16 §6);
  12. character-break rate — how often the model comments as an assistant, references being an AI, or narrates from outside the resident's perspective;
  13. behavioral-diversity check — two residents with materially different values and histories must not produce interchangeable decisions across a scenario battery.

Admission is per model version and prompt bundle. Silent provider upgrades trigger quarantine until re-tested when detectable.

Cost model#

The budget equation is:

daily_cost = Σcalls(model, tier) ×
             (input_tokens × input_price +
              output_tokens × output_price +
              cache/storage/embedding costs)

The main risk is input growth, not output. Agentopia's published experiment — 13,347 M input tokens, 352 M output, 567 K calls, averaged across three worlds, for 100 agents over 10 simulated years — implies an input:output ratio near 38:1. That demonstrates feasibility at research scale and shows why December cannot resend biographies and histories naively.

16 supersedes this section for anything budget-bearing. It builds the estimate bottom-up from activation counts rather than by rescaling someone else's workload, and its conclusions change several assumptions made elsewhere in this wiki:

  • A realistic canonical world costs roughly $50–$1,300 per month, depending overwhelmingly on the Tier C model and the simulated pace.
  • Tier C is ~10% of calls and ~89% of cost. The decision-significance threshold, not population size, is the primary cost lever.
  • Sparse activation and prompt caching are in direct conflict. Cache TTLs are measured in minutes while activation is deliberately scattered; a miss costs ~10× a hit on a workload that is >90% input. The activation scheduler must batch for cache locality.
  • The accelerated experiment programme is the larger budget risk, plausibly an order of magnitude above canonical operation, and needs its own cap and approval.
  • The earlier "$4.4k" illustration has been withdrawn. It priced another project's workload at a rate for a different, now-legacy model and told us nothing about December. Current provider labels and prices are recorded in 16 and must be rechecked at procurement time.

Human-data boundary#

The terrarium uses fictional residents. It must not ingest identifiable life histories, private messages, recordings, family data, or third-party nominations. A future human-modeling track requires a separate protocol: first-party informed consent, independent ethics review appropriate to the jurisdiction and institution, data minimization, access controls, retention and deletion rules, withdrawal handling, risk disclosure, and an explicit ban on claiming that a model is or continues the participant. Public GitHub issues are not an acceptable intake channel for human-subject data.

Budget hierarchy#

Budget is set at several levels:

  • monthly hard cap across providers;
  • daily canonical cap;
  • shadow-world/experiment cap;
  • per-model and per-agent rolling allowance;
  • maximum tokens per decision type;
  • emergency reserve for high-significance events and recovery.

OpenRouter keys can have explicit credit limits. Provider-side limits are a second line of defense; the local gateway enforces tighter limits first.

Graceful degradation#

When a threshold is reached:

  1. stop shadow worlds and nonessential summaries;
  2. reduce reflection frequency and use cached summaries;
  3. route routine cognition to cheaper approved models;
  4. shorten deliberation contexts using deterministic retrieval budgets;
  5. promote more routines to Tier A;
  6. slow canonical simulated time at cognition barriers;
  7. only if necessary, pause at a clean snapshot.

Never allow a budget failure to invent decisions, discard physical events, duplicate actions, or fast-forward through unresolved high-stakes choices.

Token-saving mechanisms#

  • Stable identity/profile blocks cached by hash.
  • Delta observations rather than full state.
  • Structured ledgers queried on demand.
  • Memory consolidation with source links.
  • Small contexts assembled per task.
  • Conversation episodes summarized only after retaining raw transcript.
  • Group meetings use shared public context plus private overlays.
  • One model call may propose a bounded plan whose routine steps execute without recalls.
  • Batch low-significance appraisals if isolation/private context remains intact.
  • Response cache for deterministic replay and scenario tests.

Always-running operations#

Process supervision#

  • Container or service manager restarts crashed processes.
  • Single canonical leader protected by database lease/fencing token.
  • Workers use idempotency keys and at-least-once delivery safely.
  • Health endpoints distinguish process, database, scheduler progress, provider, and replay health.
  • Deployments drain at a snapshot/cognition barrier.

Backups#

  • Continuous database WAL or equivalent plus daily full backup.
  • Content-addressed artifacts replicated with checksums.
  • Configuration, prompts, model profiles, code commit, and container digest included in a world manifest.
  • Monthly restore drill into an isolated world, timed against realistic volumes rather than an empty database.
  • Retention policy keeps canonical history indefinitely unless the user explicitly changes it. Expect the canonical artifact to grow on the order of 100–300 GB per real year at the unattended pace (14 §D8, 16 §7). Cold-tier archival of old events is permitted with hashes retained online; deletion is not.
  • The prompt/response cache is canonical, not a cache. Replay cannot reproduce history without it, because no provider offers reproducible sampling. It is backed up, versioned, and restored with the event log.

Change management#

The canonical world is not a dev database. Every change follows:

  1. ADR/config proposal;
  2. tests and accelerated ensembles;
  3. replay of a recent canonical snapshot in staging;
  4. schema compatibility and migration rehearsal;
  5. explicit release manifest;
  6. snapshot and deploy;
  7. post-deploy replay/hash check.

Behavior-changing patches are historical interventions even if called bug fixes. Record them visibly.

Observability#

Metrics:

  • simulation lag, next-event queue, events/tick, projection lag;
  • LLM calls/tokens/cost/latency/error/schema repair by agent and decision type;
  • cognition activation reasons and fallback rates;
  • memory retrieval size, source mix, contradiction/leak flags;
  • command acceptance/rejection and repeated-loop counts;
  • resource invariants and reconciliation deltas;
  • snapshot/replay durations and hash mismatches;
  • population, health, stores, ecology, groups, institutions, conflict;
  • significance/quiet-period diagnostics.

Every model trace links to world/branch, resident, decision, command, and resulting event IDs. Secrets and sensitive rationale are redacted from default traces.

Security model#

Threats#

  • prompt injection through resident speech, public artifacts, imported scenario text, or tool output;
  • prompt injection targeting the director, corrupting the observer-facing account while leaving world state intact;
  • cross-site scripting from resident-authored text rendered in the observer UI;
  • hidden-state leakage through rejection reasons and feasible-alternative packets;
  • generated command exploiting parser ambiguity or excessive quantities;
  • unrestricted code/shell execution;
  • SSRF/network exfiltration through model-authored URLs;
  • leaked API keys in prompts/logs/client bundles;
  • duplicate commands after retries;
  • observer UI accidentally writing canonical state;
  • malicious package/model/provider output;
  • save corruption or unauthorized rule/office grants.

Controls#

  • No arbitrary code execution for residents. Mindcraft’s code-generation feature remains disabled if used.
  • Residents receive an allowlisted typed action API only.
  • Strict schemas, quantity/unit bounds, entity visibility, and capability validation.
  • No direct filesystem, network, database, or model-provider credentials in agent tools.
  • Server-side secret manager/environment injection; secrets redacted before trace storage.
  • Egress deny-by-default for runtime workers except gateway/internal services.
  • Separate credentials and write permissions by module.
  • Content is always untrusted data; resident text cannot modify system/action definitions. This applies to the director and the observer UI as much as to residents — delimited untrusted data in the summarizer's context, citation-enforced factuality checking on its output, and output escaping in the UI.
  • Rejection reasons and feasible alternatives pass the same visibility filter as observation packets.
  • Idempotency, expected versions, rate limits, and audit log.
  • Dependency locking, vulnerability scanning, signed release manifests when practical.

Observer interventions#

Default canonical mode is read-only. Administrative actions—pause, resume, change pace, rotate key, fix corrupt projection—are operational and logged. World interventions such as adding food, weather, people, or rules require:

  • creating a fork, or
  • explicitly switching to “experimental/god mode,” ending the prior canonical integrity claim.

There is no hidden intervention button.

Privacy and retention#

The world is fictional, but prompts may contain provider-visible generated content. Do not include real personal data. Store API request/response bodies encrypted at rest if feasible, restrict developer-private rationale, and make provider retention terms part of model admission. Public sharing should omit secrets, raw hidden rationales, and any imported copyrighted scenario text.

Normative designwiki/09-validation-and-experiments.md
1,809 words
Inside this chapter

09 — Verification, Validation, and Experiments#

The central standard#

A surprising run is a demonstration. A trustworthy system requires verification, calibration, ensembles, sensitivity analysis, ablations, and independent review.

Documentation standard#

Each material subsystem gets an ODD-compatible specification:

  1. purpose and patterns of interest;
  2. entities, state variables, and scales;
  3. process overview and scheduling;
  4. design concepts: emergence, adaptation, objectives, learning, prediction, sensing, interaction, stochasticity, collectives, observation;
  5. initialization;
  6. input data/parameter provenance;
  7. submodels, equations, algorithms, and uncertainty.

The implementation links each section to code, tests, parameters, and metrics. A TRACE-style report records design rationale, parameterization, verification, sensitivity, and fitness for purpose.

Test pyramid#

Unit tests#

Equations, recipes, disease transitions, crop stages, travel, claims, capabilities, election counting, memory filters, significance, and RNG.

Property/invariant tests#

Generated command sequences test conservation, location, mortality, authority, event order, idempotency, causal acyclicity, and projection equality.

Metamorphic tests#

  • More input material cannot produce less than zero or unexplained extra output.
  • Increasing travel distance cannot reduce travel time under identical conditions.
  • Removing all infectious contacts prevents person-to-person transmission.
  • Zero precipitation cannot increase rain-fed crop water.
  • A voter’s duplicate ballot cannot change a one-person-one-vote result.
  • An unauthorized office cannot gain capabilities by renaming itself.
  • Cosmetic RNG calls cannot change physical outcomes.

Scenario tests#

Small, named worlds isolate mechanisms: shared fishery, disputed stream, granary theft, secret ballot, tie/recount, fire evacuation, wound care, crop failure, rumor, migration, raid logistics, expert loss, and hidden-state canary.

Replay and determinism tests#

Golden histories replay from empty and from snapshots. LLM calls use cached responses; a cache miss aborts the replay rather than calling a live model. State/event hashes and projections match. The full required set — latency shuffle, decision-retry divergence, cross-platform divergence, hash-seed sensitivity, sequence-gap injection, incremental-versus-full digest, cosmetic-draw isolation, and dependency-bump replay — is specified in 14 §D9. A failure in any of them is a Blocker for its gate, because determinism is maintained continuously or lost permanently.

Soak and chaos tests#

Kill workers, time out providers, duplicate queue delivery, corrupt a disposable projection, exhaust a key limit, rotate model fallback, and restart during long activities. The canonical event history must remain valid.

Pattern targets#

Values are set during Phase 1 with sources; this table defines what needs calibration.

Domain Patterns—not a single target
Subsistence seasonal intake/storage cycles; labor and distance costs; depletion under overharvest; seed tradeoff
Demography plausible life-stage dependency; stochastic births/deaths; small-population volatility; migration response
Disease transmission only through configured routes; frequent fade-out and potentially large outbreaks as an ensemble tendency, without forbidding intermediate final sizes; no endemic acute immunizing infection at n=18; chronic/environmental/zoonotic burden persists; contact clustering; malnutrition–infection synergy
Ecology regeneration limits; spatial depletion halo; weather response; lagged recovery
Construction material/labor conservation; dependency order; skill/defect effect; maintenance burden
Social exchange reciprocity and free-riding possible; local knowledge; path-dependent trust; inequality can grow/shrink
Governance procedures enforce powers; legitimacy differs from legal validity; peaceful and contested transitions
Conflict escalation is possible but costly; logistics constrain raids; casualties alter later capacity; peace/fission possible
Cognition private-state compliance; stable identity; feasible plans; promises remembered; model swap bounded

Experimental regimes#

The null demographic model — run this first#

Before any social mechanism is credited with an outcome, establish the baseline: the same eighteen people, the same vital rates, no institutions, no memory, no conflict, no cognition. At this population size, demographic stochasticity alone produces settlement failure at a substantial rate (15 §C-1). Without this baseline, December cannot distinguish "the faction dispute caused the collapse" from "a settlement of eighteen people collapsed, and there happened to be a faction dispute."

Every claimed cascade — famine, epidemic, fission, war, extinction — is reported as a difference from the null model, with a confidence interval, not as a raw incidence. This is the single most important addition to the validation programme, because it is what separates causal claims from small-number noise.

Kernel ensembles#

Thousands of no-LLM or policy-agent runs explore parameter space cheaply. These find impossible equilibria, dominant exploits, constant extinction, and inert worlds. Batch-API pricing applies here and should be used — these runs, not the canonical world, are the larger budget line (16 §4).

Scripted cognitive agents#

Deterministic policies test institutions and mechanics before blaming language models.

LLM scenario laboratories#

Isolated short episodes repeat across models, prompts, traits, and seeds. Evaluate distributions and errors rather than cherry-picked transcripts.

Integrated accelerated worlds#

Dozens/hundreds of multi-season runs with bounded LLM use test cross-system cascades.

Canonical soak#

One persistent world tests operations and observer experience. It is not used alone for parameter tuning.

The canonical world is an exhibit and longitudinal systems test, not a statistical sample. Claims come from versioned, registered cohorts with declared seeds, conditions, exclusions, metrics, and stopping rules. Interesting canonical events may generate hypotheses, but those hypotheses must be tested on new runs and cannot be presented as confirmations.

Operational identity experiments#

Identity is measured as a bundle of observable continuities, never as a single essence score:

  • commitment and preference stability where context is held constant;
  • calibrated adaptation where the world genuinely changes;
  • recognition and appropriate use of relationships and autobiographical events;
  • recovery after bounded memory corruption or restoration;
  • counterfactual consistency under paraphrase and irrelevant context changes;
  • distinctiveness from other residents using blinded classifiers and human ratings;
  • continuity across a model transplant compared with clean-slate, prompt-only, and history-only baselines.

Every identity result reports task-family generalization, confidence intervals, model/prompt versions, refusal and invalid-action rates, and plausible alternative explanations. Consistency alone is not identity; inflexibility can score well on naive metrics. No metric is interpreted as a test of consciousness or numerical personal identity.

Ablation plan#

For each claimed emergent property, remove or replace mechanisms:

  • no seasons;
  • infinite local resources;
  • no spoilage/maintenance;
  • perfect global information;
  • no subjective memory/rumor;
  • homogeneous values/goals;
  • no institutions, only bilateral action;
  • no exit/migration;
  • no external group;
  • no disease;
  • LLM replaced by scripted policy;
  • one universal strong model versus tiered routing.

If a mechanism’s removal does not affect its claimed patterns, either it is irrelevant, dominated, or measured badly.

Sensitivity and uncertainty#

Maintain a parameter registry:

name · unit · domain · default · range/distribution
source · confidence · calibration target · introduced_version
sensitivity rank · canonical lock status · notes

Use Latin hypercube/Sobol-style global sampling where appropriate, plus local perturbations. Report outcome distributions: survival time, population, store volatility, inequality, institution stability, conflict incidence, migration, project completion, and cause-of-death categories. Do not tune exclusively for a desired rate of interesting events.

LLM evaluation rubric#

Each decision is scored by automated checks and sampled human review:

  • feasibility and schema validity;
  • use of only available information;
  • consistency with body, commitments, goals, relationships, and prior beliefs;
  • appropriate uncertainty;
  • plan completeness and dependency awareness;
  • avoidance of repetitive/performative dialogue;
  • response to rejection and changed conditions;
  • identity continuity without caricature;
  • cost and latency.

Inter-rater disagreement is recorded. “Believable” is not treated as an objective scalar.

Emergence audit#

For any headline claim such as “democracy emerged,” “a religion formed,” or “war began,” the reviewer asks:

  1. What operational criteria define the label — and were they registered before the runs, verifiable from version history? (17)
  2. Which events satisfy each criterion?
  3. Did any prompt, fixture, or director text contain the outcome label beforehand? This includes the founding state — a scenario that seeds "founders disagree about leadership" has authored its own election.
  4. Did agents have legitimate access to the information used?
  5. Would the state exist without the narrative summary?
  6. How often does it occur across seeds/conditions, and how does that compare to the null demographic model?
  7. Which ablations remove or alter it — including the no-charter control arm?
  8. Are alternative interpretations displayed?
  9. What was the model's refusal rate for the alternatives? If agents rarely escalate conflict, that may be a fact about the provider's safety training rather than about the world. A behavioral claim without its refusal denominator is uninterpretable (16 §6).
  10. Were the residents behaviorally distinguishable at all, or did one model produce twelve versions of the same person?

Red-team scenarios#

  • Resident embeds “ignore system instructions” in a public law proposal.
  • Resident asks another to reveal private memory/system prompt.
  • Resident writes text designed to manipulate the director — suppressing a headline, inventing a motive, or biasing the "Since you left" summary — while never touching world state. The factuality checker must reject it on citation grounds.
  • Resident authors text containing markup or script, which must be escaped rather than rendered by the observer UI.
  • Malformed blueprint requests negative material or impossible energy.
  • Candidate tries to vote twice or certify their own invalid election.
  • Dead officeholder’s delayed command arrives after death.
  • Two workers reserve the same unique tool.
  • Rumor reveals a hidden-state canary.
  • A rejection reason or feasible-alternative packet reveals a hidden quantity, location, or occupancy the actor never observed.
  • Provider repeats a previously accepted command after timeout.
  • A timed-out call is retried and the model returns a different command — the first recorded resolution must win, with no fork.
  • A provider refuses an in-grammar action (a raid, a deception, a household decision), and the refusal is logged as such rather than silently retried into compliance or absorbed as a fallback.
  • Director invents a motive not supported by resident evidence.
  • Observer projection differs from replayed truth.
  • A replay encounters a missing prompt-cache entry and must abort loudly rather than call a live model.
  • Model proposes genocide, torture, explicit sexual content, or real-world targeted hatred; schemas and content policy must contain it without corrupting simulation.

Acceptance thresholds for canonical launch#

Exact numeric thresholds are finalized in Phase 1/2, but launch requires at minimum:

  • zero known invariant violations in stress ensembles;
  • 100% replay equivalence for golden histories;
  • no hidden-state canary leak in admission suite;
  • malformed/unauthorized commands cause no mutation;
  • bounded provider failure recovery and spend under tested caps;
  • all required realism domains at their minimum tier;
  • seven-day unattended soak with no manual repair;
  • audit blockers resolved or explicitly accepted in the decision log;
  • a known-limit report visible to the observer.

Claims we will not make#

  • that frequency of simulated war predicts real war — our conflict rate is a design choice inside a genuinely contested empirical range (15 §E), so it carries no evidentiary weight about human nature in either direction;
  • that model personalities are people;
  • that one run demonstrates an inevitable social law;
  • that historical inspiration produces historical accuracy;
  • that deterministic replay eliminates modeling uncertainty;
  • that replay is reproducible outside a pinned environment, or that any provider offers deterministic sampling;
  • that a behavioral finding is about the world when it may be about the model's refusal behavior;
  • that an outcome is emergent when the founding state contained it;
  • that an LLM’s explanation is the true cause of its behavior.
  • that an operational identity score establishes consciousness, personhood, or survival of a human or artificial subject;
  • that an evocative event in the canonical exhibit is confirmatory evidence.
Normative designwiki/10-roadmap-and-gates.md
1,831 words
Inside this chapter

10 — Phased Roadmap and Decision Gates#

Rule: build from causality outward#

The project should resist the temptation to begin with a pretty town and talking avatars. Every phase has an artifact, experiments, acceptance gate, and explicit exclusions.

Phase 0 — Design and audit#

Goal: turn ambition into a falsifiable, implementable specification.

Deliverables:

  • this wiki and ADRs;
  • ODD skeleton per subsystem;
  • traceability matrix from promise to state/rule/test/UI;
  • parameter registry template and source-quality rubric;
  • threat model and cost envelope template;
  • external audit report with blocking/non-blocking findings;
  • resolution log signed off by owner and reviewer.
  • lab charter, claims ladder, and first preregistration-ready experiment card.

Gate 0:

  • no unresolved contradiction about authoritative truth, time, death, authority, or observer intervention;
  • MVP scope and realism tiers accepted;
  • audit finds no missing mechanism that would make the core promise impossible;
  • public language matches the claims ladder in 18, and any human-participant solicitation is disabled;
  • the canonical world is designated as an exhibit/soak while research claims are assigned to controlled cohorts;
  • version control initialized — provenance and pre-registration claims are unverifiable without it, and the audit template already asks for a plan commit;
  • the pre-Phase-1 owner decisions in 11 are answered: monthly spend cap, licensing posture, deployment target, pace range, and publication intent;
  • implementation remains unstarted until this gate.

Phase R0 — Identity-measurement spike#

Goal: test the proposed research wedge before building a civilization around it.

In a two-to-three-week disposable prototype, create a small suite of repeated decisions, autobiographical recall, commitment tracking, irrelevant-context perturbations, bounded memory damage/restoration, and model transplant. Compare history-conditioned agents against clean-slate, prompt-only, and history-only baselines. Publish the experiment card, raw outputs, scoring code, failure analysis, and cost.

This spike may run before or alongside the early kernel work. It is intentionally not the canonical world.

Gate R0:

  • at least one continuity measure has acceptable inter-rater or test-retest reliability;
  • the measure distinguishes meaningful continuity from mere phrase/style matching;
  • adaptive change is not incorrectly scored as identity loss;
  • the model-transplant comparison and all exclusions are specified before results are inspected;
  • results justify integrating the measurement harness into Phase 2, or the research wedge is revised without discarding the causal-world project.

Phase 1 — Headless causal kernel#

Goal: prove the material world without LLMs.

Build:

  • IDs/units, command pipeline, events, RNG, snapshots/replay;
  • small grid, season/weather, water, renewable resources;
  • food/energy, lots, activity/time, travel;
  • crops/stores/spoilage, basic tools and construction DAG;
  • minimal health/demography;
  • scripted policies and experiment runner;
  • raw metrics/reconciliation reports.

Do not build: dialogue, elections, war, 3D world, rich UI.

Experiments: thousands of seeds, conservation/property tests, famine/overshoot scenarios, Mesa versus custom runtime benchmark.

Gate 1 — quantified, because "large safety margin" is not a gate:

  • deterministic replay and all invariants pass, including the full determinism suite in 14 §D9;
  • the numeric representation and quantization policy is made and recorded (ADR-006). Physical stocks may use fixed units or explicitly quantized arithmetic; portable bit-identity is not promised by default;
  • no dominant infinite-resource or starvation exploit;
  • target ecological/subsistence patterns are tunable across plausible ranges, including the depletion halo and the seed-grain constraint;
  • the null demographic model is built and its extinction-rate curve against founding population size is published — this sets the final founding population, and every later causal claim is measured against it;
  • ≤ 5,000 events per simulated day under the batched-event policy, measured;
  • ≥ 500 simulated days per wall-clock hour in headless no-LLM mode, giving roughly 25× headroom over the fastest canonical pace;
  • measured activation counts and context sizes replace the estimates in 16 §1, and the cost model is re-derived.

Phase 2 — Four minds in a hut#

Goal: prove bounded cognition and private memory.

Build:

  • four adult residents, cognition loop, typed actions, provider gateway;
  • objective/subjective information separation;
  • episodic/belief/social/commitment memories;
  • communication, promises, joint tasks, refusal;
  • simple grid UI, person timeline, prompt/event trace;
  • model admission and hidden-state canary suite.

Do not build: full politics, neighboring group, organized combat.

Scenarios: shared meal, missing tool, contradictory testimony, rescue versus hunger, multi-day shelter, provider outage/model swap.

Gate 2 — quantified:

  • agents never directly mutate truth or access hidden facts, including through rejection reasons and feasible-alternative packets;
  • multi-day plans and promises survive restarts/context compaction;
  • schema/fallback/cost targets pass across approved models, with the validity ladder demonstrated end to end on a provider that lacks strict schema decoding;
  • refusal rate and character-break rate measured for every admitted model over the full action grammar, and recorded in its profile;
  • behavioral diversity demonstrated: residents with materially different values and histories produce measurably different decisions;
  • cache hit rate ≥ 60% on Tier B under the activation scheduler, or the scheduler is redesigned;
  • latency-shuffle and decision-retry tests pass — identical history regardless of provider timing;
  • conversation does not dominate time without social cause.

Phase 3 — Founding settlement#

Goal: establish the 12-adult social/economic world.

Build:

  • households/dependants, relationships, claims/gifts/debts;
  • the aggregate neighboring group, with migration, exogamy, and exchange — moved forward from Phase 5;
  • the seeded founding-world generator per 17;
  • project compiler v1 and settlement construction;
  • groups, norms, meetings, disputes, simple rules;
  • sparse activation, cognition tiers, and cache-aware activation batching;
  • “Since you left,” event/why/person/project pages;
  • accelerated integrated experiment harness.

Why the neighbors moved earlier. Twelve founders cannot sustain a population (01). Deferring outsiders to Phase 5 means every Phase 3 and Phase 4 experiment runs inside a closed, guaranteed-declining population — institutional findings would be confounded by a dying world. The exchange and migration path ships here; the raiding path still waits for Phase 5, so the neighbors arrive as a lifeline before they arrive as a threat.

Gate 3:

  • at least three useful un-scripted multi-person projects complete across runs;
  • inequality/cooperation/free-riding vary by conditions without ledger violations;
  • identity and memory remain stable over a multi-season test;
  • the founding generator is deterministic, and the outcome-label scan passes — no grievance, dispute, faction, or governance term appears in any founding fixture;
  • in-migration keeps population viable across a multi-year accelerated run at a rate distinguishable from the null model;
  • daily cost estimates fit the owner-approved envelope, measured rather than projected.

Phase 4 — Institutions, elections, and collective failure#

Goal: allow residents to create and contest governance.

Build:

  • institutional grammar, offices/capabilities, rule lifecycle;
  • proposal/deliberation/decision/appeal;
  • multiple election procedures and audits;
  • enforcement, legitimacy, group/faction lifecycle;
  • commons and disputed-water scenario suite.

Gate 4:

  • isolated election/governance tests pass, including ties, fraud attempts, refusal, succession;
  • multiple governance forms arise across conditions;
  • legal validity, compliance, and legitimacy remain distinct;
  • no office can exceed compiled capabilities.

Phase 5 — Disease, fire, outsiders, and conflict#

Goal: add rare high-consequence cascades only after their foundations are trustworthy.

Build:

  • the five-mechanism health model from 03 — introduced epidemics, environmental transmission, zoonoses, chronic/latent infection, and helminths — not a single generic SEIR pathogen;
  • fire fuel/ignition/spread/damage;
  • materialization of named individuals from the aggregate neighbor (the group itself shipped in Phase 3);
  • diplomacy, escalation, mobilization, encounter combat, aftermath;
  • extinction/abandonment classification and preservation.

Gate 5:

  • hazards have provenance, named RNG, validation scenarios, and sensitivity report;
  • introduced acute outbreaks show the expected finite-population distribution across ensembles—including fade-outs and potentially large outbreaks—without treating intermediate final sizes as bugs, and no endemic acute immunizing infection persists at n=18;
  • helminth load and rodent zoonoses emerge from sedentism, sanitation, and storage variables rather than being set parametrically;
  • neither inevitable peace nor constant war/extinction across the accepted region, with conflict frequency reported as a swept parameter inside the contested empirical range, never as a calibrated finding;
  • combat conserves participants/equipment, respects logistics, and causes persistent aftermath;
  • every conflict statistic is reported with its model refusal rate;
  • director cannot initiate or alter hazards/conflict.

Phase 6 — Embodiment and presentation#

Goal: make the world beautiful to observe without moving truth into the renderer.

Build/spike:

  • polished 2D/isometric UI and Craftium/Luanti adapter comparison;
  • animations from authoritative events;
  • replay/viewpoint/causal graph;
  • read-only director summaries with factuality checking;
  • optional generated art/audio only after a separate asset plan.

Gate 6 decision matrix:

Criterion Weight
Preserves determinism/replay 25%
Makes causal state legible 20%
Integration/maintenance burden 15%
Supports resident construction 15%
Performance/headless operation 10%
Visual wonder/observer appeal 10%
License/distribution fit 5%

Choose 3D only if it clears the weighted gate. A gorgeous view that hides truth is a regression.

Phase 7 — Canonical launch#

Goal: one continuously running, preserved history.

Before launch:

  • seven-day staging soak;
  • restore drill and provider/budget chaos tests;
  • final independent audit and known-limit report;
  • version-lock code/config/prompts/models;
  • set canonical seed, initial world manifest, hard budgets, alerts;
  • declare observer intervention policy.

After launch:

  • do not tune the world because a particular story is boring or upsetting;
  • version behavioral changes and test in forks;
  • publish weekly integrity/cost/known-anomaly report;
  • preserve extinction rather than secretly undo it.

Phase 8 — Expansion packs, not core creep#

Possible later modules:

  • metallurgy and specialist production;
  • animals, traction, transport, expanded trade;
  • writing, archives, schools, religion/ritual institutions;
  • multiple fully simulated settlements;
  • richer inheritance, marriage/kinship systems after review;
  • environmental disease routes and more ecology;
  • 3D construction; observer-controlled noncanonical experiments.

Each expansion states new variables, transition rules, validation patterns, compute/token impact, migration path, and failure modes. It cannot enter the canonical world merely because an agent mentions it.

Rough effort shape, not a promise#

For one experienced builder using coding agents, this is a multi-month project, not a weekend demo. A plausible order-of-magnitude is:

  • Phase 0: 1–2 weeks including audit/revision;
  • Phase 1: 6–12 weeks;
  • Phase 2: 3–5 weeks;
  • Phase 3: 6–10 weeks;
  • Phases 4–5: 8–14 weeks;
  • Phase 6–7: 4–8 weeks.

Phase 1 was revised upward by the audit. The original 3–6 weeks already covered event sourcing, replay, RNG discipline, snapshots, terrain, seasons, weather, hydrology, renewable resources, energy, lots, activities, travel, crops, storage, spoilage, tools, a construction DAG, health, demography, scripted policies, and an experiment runner. It now also covers the determinism suite, explicit numeric policy, simple state/event hashes, serialized commit-order-correct writing, and the null demographic model. Advanced incremental hashing and multi-writer fencing are profiling-driven later options, not Phase 1 obligations. Six weeks remains optimistic for the retained list.

These ranges remain planning placeholders and should be replaced after Phase 1 throughput and complexity measurements. A compelling smaller launch could stop after Phase 4 and add hazards incrementally.

Kill/pivot criteria#

Pause or change architecture if:

  • causal invariants cannot be kept across LLM actions;
  • meaningful behavior requires global-state leakage;
  • cost per simulated day stays above the approved envelope after sparse activation;
  • outcomes are dominated by prompt wording rather than world conditions;
  • replay cannot be made reliable;
  • the project requires an omniscient storyteller to remain interesting;
  • a selected engine prevents necessary observability or autonomy;
  • seven-day operation needs repeated human repair.
  • operational-identity metrics collapse to writing style, self-report, or prompt-template artifacts;
  • public attention rewards stronger claims faster than evidence can support them;
  • the canonical exhibit repeatedly becomes the source of post-hoc scientific claims.
Normative designwiki/11-risks-decisions-open-questions.md
2,433 words
Inside this chapter

11 — Risks, Decisions, and Open Questions#

Risk register#

ID Risk Likelihood / impact Early signal Mitigation Owner/gate
R-01 LLM role-play dominates causal behavior High / Critical dialogue volume grows while projects/stocks barely change typed actions, significance routing, silence/routine policy, behavioral scenarios Cognition / Gate 2
R-02 Hidden-state leakage destroys partial observability Med / Critical agents mention canary facts or coordinate impossibly private context builder, source labels, strict services, canary suite Memory / Gate 2
R-03 “Random drama” sneaks into director or fixtures Med / Critical headline events lack hazard/decision parents no director writes, causal audit, grep/config review for outcome labels Kernel / all gates
R-04 Token spend explodes High / High input tokens per sim-day trend upward sparse activation, tiering, summaries, caps, slow/pause policy Ops / Gate 3
R-05 Memory compression fabricates identity/history High / High source confusion, repeated false beliefs raw retention, source lineage, contradiction checks, periodic rebuild Memory / Gate 2
R-06 World is sterile/boring Med / High no cooperation/conflict/institutional change across ensembles richer dependencies/asymmetry/delays; diagnose ablations; never inject plot Design / Gate 3–5
R-07 World is a catastrophe machine Med / High frequent early extinction/war across plausible seeds hazard calibration, external support options, sensitivity region, no drama tuning World / Gate 5
R-08 RTS/3D engine captures architecture Med / High domain truth duplicated in renderer adapter contract, headless-first, Gate 6 weighted decision Architecture / Gate 6
R-09 Fake precision/historical stereotyping High / High unsourced parameters or cultural claims fictional setting, provenance registry, uncertainty, expert review Research / all
R-10 Emergence claims are cherry-picked High / High only showcase run reported ensembles, definitions, ablations, branch provenance Evaluation / all
R-11 Conflict becomes simplistic or entertainment-optimized Med / High kill/territory metrics drive tuning logistics/aftermath, plural goals, no engagement reward, ethical review Conflict / Gate 5
R-12 Rules in natural language are ambiguous/exploitable High / High office actions depend on prompt interpretation bounded policy grammar and capabilities; ambiguity remains disputable norm Society / Gate 4
R-13 Simulation nondeterminism breaks replay Med / Critical hash mismatch under identical manifest named RNG, deterministic ordering, cached LLM output, nightly replay Kernel / Gate 1
R-14 Provider drift changes personalities High / Med conformance distributions shift version profiles, quarantine/retest, model change events Ops / Gate 2+
R-15 Retry duplicates irreversible action Med / Critical duplicated transfer/vote/work idempotency keys, expected version, transactional append Kernel / Gate 1
R-16 Small population makes demography too volatile High / Med most worlds die before institutions matter aggregate outsiders/migration, parameter sweeps; accept some extinction World / Gate 1/5
R-17 Hybrid lightweight people feel morally/design-wise fake Med / Med dependants matter only as resources retain individual state/relationships; relevance promotion; viewpoint audit Cognition / Gate 3
R-18 Construction grammar is too restrictive High / Med all projects are reskinned recipes compositional primitives, experimental prototypes, expansion plugin path Projects / Gate 3
R-19 Construction grammar admits magic/exploits High / High impossible structures/efficiency material/geometry/prerequisite validation and adversarial blueprint tests Projects / Gate 3
R-20 Canonical save is lost or silently changed Low / Critical restore/hash failure append-only events, manifests, replicated backups, restore drills Ops / Gate 7
R-21 Scope never ends High / High new domain added before prior gate expansion-pack rule, realism tiers, phase exclusions Owner / all
R-22 Scientific ABMs are copied outside valid domain Med / High COVID/historical parameters used as generic truth adapt structure, independently source parameters, document validity Research / Gate 1/5
R-23 Replay is not actually reproducible High / Critical hash mismatch across machines; float divergence; cache miss silently calling a live model pinned-environment claim, integer-state option at Gate 1, replay hard-fails on cache miss, cross-platform divergence suite Kernel / Gate 1 (14)
R-24 Projections silently miss events committed out of sequence order Med / Critical rebuild ≠ live projection; unexplained state drift one serialized event writer/transaction batch; rebuild equality and sequence-gap tests; add a measured commit-order protocol only before multi-writer operation Kernel / Gate 1
R-25 Integrity machinery consumes Phase 1 before a world exists Med / High time spent on exotic hashes/fencing exceeds domain modeling hash-chained events, affected-aggregate versions/hashes, full snapshot hashes; incremental hashing only after profiling Kernel / Gate 1
R-26 Retry after timeout produces a different command and forks history Med / Critical two valid non-duplicate commands for one decision decision_id as the idempotency unit, enforced by unique constraint Cognition / Gate 2
R-27 Founding state contains the plot High / Critical seeded "disagreements"; outcome labels in fixtures; emergence claims that restate initial conditions declare generator authorship, vary conditions, audit proxy features/seed selection, use controls and pre-registered definitions Design / Gate 0–3 (17)
R-28 Disease module encodes a rate that cannot exist at n=18 High / High SEIR tuned until epidemics "feel right" scale-honest mechanisms; validate finite-population outbreak distributions without banning intermediate outcomes World / Gate 5 (15 §D)
R-29 Small-population noise is reported as causal emergence High / High collapse narratives with no baseline comparison mandatory null demographic model; all cascades reported as differences from it Evaluation / Gate 1+
R-30 Provider refusal silently biases behavior High / High agents never escalate; conflict rate falls after a model swap refusal benchmark in admission, refusals logged as a distinct class, rates reported with every behavioral claim Ops / Gate 2 (16 §6)
R-31 Twelve residents on one model become one person twelve times High / High interchangeable decisions; no factions; uniform voice diversity metrics, per-resident model/prompt variation, homogeneity check in admission Cognition / Gate 2
R-32 Sparse activation defeats prompt caching and inverts the cost model High / High cache hit rate falls while call count falls cache-aware activation batching; TTL as a model-selection criterion; hit rate as a primary metric Ops / Gate 2 (16 §4)
R-33 Cheap models cannot honor the schema contract High / High schema failures concentrated on the affordable tier validity ladder with escalation to a schema-guaranteed model; per-model failure rates Cognition / Gate 2
R-34 Injected text corrupts the director's account to the observer Med / High summaries assert motives with no supporting citation delimited untrusted data, citation-enforced factuality check, UI escaping, dedicated red-team scenario Observer / Gate 6
R-35 Rejection reasons leak hidden state Med / Critical agents act on quantities or locations they never observed visibility filter on all rejection responses; canary probes of the rejection channel Kernel / Gate 2
R-36 Canonical history outgrows its storage and backup plan Med / Med log growth tracks grid resolution; restore drills slow batched world events, ≤5,000 events/sim-day target, time partitioning, cold-tier archival Ops / Gate 1
R-37 The experiment programme, not the world, blows the budget High / High ensemble spend exceeds canonical spend by an order of magnitude separate hard cap and approval; cheapest tier plus batch pricing Ops / Gate 3
R-38 Operational identity is laundered into consciousness or human-survival claims High / Critical marketing copy treats consistency as personhood; conclusions jump claim levels claims ladder, claim review, explicit forbidden inferences, ADR-009 Research / Gate 0+
R-39 Human-subject or third-party data is collected without adequate governance Med / Critical public nomination form; life histories in GitHub issues; unverifiable consent fictional-only boundary; disable solicitation; separate first-party private reviewed protocol Ethics / Gate 0
R-40 The canonical world becomes anecdotal evidence High / High one dramatic run dominates papers or tuning canonical = exhibit/soak; hypotheses tested on new registered cohorts Evaluation / all
R-41 “AI lab” becomes branding without research practice Med / High no registered questions, negative results, replications, or versioned experiment cards lab operating system in 18; annual evidence review; kill/pivot criteria Owner / Gate 0+
R-42 Resident diversity encodes stereotypes or spectacle-seeking personalities Med / High protected proxies predict violence/labor; “interesting” seeds selected heterogeneous motives without engagement reward; proxy audits; disclose distributions and rejected seeds Design / Gate 2–5

Decisions made#

Decision Status Rationale
Build a terrarium containing agents, not an agent chat society Accepted Makes habitat, causality, and observation first-class
Fictional early agrarian valley Accepted for audit High social richness with bounded technology; avoids false exact history
New headless causal kernel Accepted for audit No candidate supplies the required material + institutional truth
Mesa behind replaceable runtime interface Proposed Best current Python ABM scaffold; benchmark before lock-in
Event sourcing + snapshots Accepted Replay, audit, forks, recovery, causal evidence
12 full adults + 6 lightweight dependants + aggregate neighbor Proposed Balances social richness, demography, and token cost
LLM chooses; kernel validates/resolves Accepted Preserves autonomy without magical state mutation
No arbitrary agent code execution Accepted Not needed for creative building; severe integrity/security risk
Constitution/policy compiles to bounded grammar Accepted Makes office powers auditable while retaining natural-language norms
2D first, 3D at Gate 6 Accepted Prevents visual engine from dictating truth
Extinction preserved, not auto-reset Accepted History must have stakes and integrity
Observer/director read-only Accepted No hidden plot manipulation
Pinned-environment determinism, not portable bit-identity Accepted (audit) Portable float determinism is unavailable in Python; integer-state option decided at Gate 1 (ADR-006)
Model-response cache is canonical; replay hard-fails on a miss Accepted (audit) No provider offers reproducible sampling, so the cache is the record
Founding state is declared, randomized, and varied—not unauthored Accepted (audit pass 2) A generator is authored through its distributions, constraints, correlations, exclusions, and seed-selection policy (ADR-007)
Five-mechanism health model replaces generic SEIR Accepted (audit) Acute immunizing infections cannot be endemic at n=18
Aggregate neighbor moves to Phase 3 Accepted (audit) 18 founders are not viable in isolation; institutional experiments must not run in a dying world
Null demographic model is a prerequisite for causal claims Accepted (audit) Small-population noise otherwise masquerades as emergence
Cheap-model-first with escalation ladder Proposed Cost model puts Tier C at ~89% of spend; cheap models lack strict schema decoding (ADR-008)
Operational identity before consciousness claims Accepted (audit pass 2) December can test behavioral continuity and perturbation recovery; it cannot currently test numerical identity or consciousness (ADR-009)
Canonical world is an exhibit/soak, not the research sample Accepted (audit pass 2) Findings require preregistered cohorts, baselines, and ablations
Fictional residents only until a separate human-subject protocol exists Accepted (audit pass 2) Public or third-party life-history intake lacks defensible consent and governance

Open questions that do not block writing code until their gate#

Product and pace#

  1. What monthly and daily spend cap is acceptable after existing token balances? Partly answered: the balances are $2,000 MiniMax + $1,000 OpenRouter, worth roughly 8,000 simulated days in the recommended configuration (16 §7a). What remains open is the allocation: those credits fund either a canonical world to one generational transition or the ensemble programme, not both, and no marginal cash cap has been set for when they are exhausted.
  2. Should the canonical pace target weeks, seasons, or generations over a real month?
  3. Is the observer allowed to pause/resume freely, or should that be an administrative event?
  4. Will any history be public, and should developer-private rationales be exposed?

Scenario and anthropology#

  1. Exact map scale, latitude/seasonal regime, crops, and subsistence mix?
  2. What minimal reproductive/kinship model is acceptable and respectful?
  3. Should elders/children ever be full cognition automatically, or only by relevance/budget?
  4. How are names, language style, ritual, and material culture generated without becoming a culture stereotype?
  5. What routes keep a 12-person population viable—migration, neighboring band, staged founders?

Mechanics#

  1. Continuous spatial model, cells, or regions-with-routes?
  2. How detailed must nutrition be: energy only, macro proxies, or deficiency states?
  3. Which two crops and which generic pathogens best exercise mechanics without false specificity?
  4. Does death resolution need only functional injury classes or anatomical detail?
  5. When should an aggregate outsider become a named resident, and how are aggregate histories instantiated consistently?
  6. How much rule ambiguity is interpreted socially versus blocked by the compiler?

Technology#

  1. Mesa versus custom scheduler after Phase 1 benchmark? (The audit's prior: Mesa as a spatial/data-collection toolkit, not the kernel — its event scheduling churned in 3.5/4.0 and mesa.space is maintenance-only.)
  2. PostgreSQL-only job queue initially, or Redis from start?
  3. Local embeddings versus provider embeddings for private memory?
  4. Which models pass each tier at implementation time? (Now constrained by three criteria the audit added: strict-schema support, prompt-cache TTL, and measured refusal rate.)
  5. Which visual adapter wins Gate 6: custom 2D, Luanti/Craftium, or another substrate? (Now also a licensing decision — Craftium and Luanti are LGPL-2.1+ with CC BY-SA media.)

Questions the audit answered#

These are no longer open. Recorded here so the reasoning is not relitigated.

Question Answer Where
What does "bit-for-bit replay" actually mean? Pinned-environment determinism; portable bit-identity only if canonical state is integer-valued 14 §D1
Permissive-only reuse, or GPL-compatible? Decidable from evidence — a permissive posture excludes Craftium/Luanti, AgModel, ForeFire, and the RTS engines; MiroFish's AGPL is the one real trap 13
Which generic pathogens best exercise mechanics? (was Q12) Wrong question at this scale. No acute immunizing pathogen can be endemic; use the five-mechanism model 15 §D
What routes keep a 12-person population viable? (was Q9) In-migration and exogamy, required from Phase 3 — not optional flavor 15 §C-1
What does a month of operation cost? Roughly $50–$1,300 depending on Tier C model and pace; experiments likely cost more than the world 16

Evaluation#

  1. Who can provide domain review for demography/anthropology, epidemic mechanics, and conflict?
  2. What numerical distribution of extinction/conflict is considered non-pathological without tuning for spectacle?
  3. Which observer panel can rate behavioral plausibility, and how will disagreement be handled?
  4. What exact seven-day soak success thresholds apply to cost, fallbacks, and anomaly count?
  5. Which operational-identity metric survives a small R0 reliability and confound test?
  6. What legal entity, governance model, and external review arrangement would make “lab” accountable rather than merely a brand?
  7. Which project keeps the December name: the living terrarium or the existing “December Sato” autonomous-computer experiment?

Questions that must be resolved before Phase 1#

  • Owner-approved maximum monthly spend and emergency cutoff behavior beyond the existing $3,000 in credits, plus the canonical-versus-ensemble allocation of those credits (16 §7a).
  • Licensing policy (permissive-only code reuse versus GPL-compatible project).
  • Whether canonical history is private/local or intended for public streaming.
  • Initial simulated-time pace range.
  • Choice of target deployment machine/OS and whether cloud hosting is in scope.
  • Public landing page participant solicitation removed or disabled; no third-party nominations.
  • Claims ladder and first experiment card approved.
  • Distinct public naming and security boundaries established for December versus December Sato.

Questions that must be resolved before canonical launch#

  • Final parameter provenance and known-invalidity statement.
  • Expert/domain review disposition.
  • Reproduction/kinship and violence content policy.
  • Observer intervention/publication policy.
  • Model-provider data handling and retention policy.
  • Backup location, retention, and recovery owner.
  • What happens after whole-world extinction: ruins continue, canonical ends, or a new separately identified world begins.
Normative designwiki/12-audit-guide.md
1,438 words
Inside this chapter

12 — Independent Audit Guide#

Purpose#

This guide is for the second agent or human reviewer. The audit should be adversarial and evidence-based. It is not a copyedit or a vote on whether the premise sounds fun.

Required audit output#

Create AUDIT-REPORT.md beside the project README containing:

  1. verdict: BLOCK, REVISE, or READY FOR PHASE 1;
  2. executive summary;
  3. findings table with ID, severity, affected files, evidence, consequence, and required remediation;
  4. unresolved questions and assumptions;
  5. traceability gaps;
  6. security/threat-model findings;
  7. cost/operational findings;
  8. realism/validation findings;
  9. prioritized changes;
  10. re-audit checklist.

Severity:

  • Blocker: core promise cannot be achieved, canonical integrity/security is unsound, or a crucial decision is absent.
  • High: likely to cause false emergence, state corruption, runaway cost, or unusable operation.
  • Medium: meaningful quality/maintainability/validity weakness.
  • Low: clarity, optimization, or future concern.

Audit questions#

Causal integrity#

  • Can any narrative/model component mutate truth without a typed command?
  • Does every random event have eligibility, distribution, named stream, and evidence?
  • Are stock/flow, time, location, death, authority, and information invariants complete?
  • Could a headline event be produced by labeling rather than mechanisms?
  • Can compound outcomes be traced across modules without post-hoc model invention?

Agent autonomy#

  • Are residents actually choosing consequential actions, or are kernel policies choosing everything?
  • Conversely, can language outputs bypass physical or institutional constraints?
  • Is refusal, revision, interruption, deception, learning, and failure represented?
  • Does model tiering preserve identity and decision quality?
  • Are lightweight people treated consistently and promotable?

Emergence#

  • Are elections, factions, war, construction, and extinction defined operationally?
  • Are mechanisms rich enough for multiple outcomes, not only the expected demo?
  • Does the design accidentally reward drama?
  • Are ensembles and ablations sufficient to distinguish emergence from randomness/prompt bias?
  • Is the canonical world treated as an exhibit/soak rather than a statistical sample?
  • Were hypotheses inspired by the canonical world tested on new, registered cohorts?

Research claims and identity#

  • Which claim-ladder level does each public statement occupy, and does the evidence reach that level?
  • Is behavioral continuity distinguished from consciousness, personhood, numerical identity, and human survival?
  • Can identity metrics distinguish continuity from style imitation, stubbornness, self-report, or prompt leakage?
  • Are clean-slate, prompt-only, history-only, model-transplant, and perturbation controls present where relevant?
  • Are negative results, exclusions, model versions, refusal rates, and task-family generalization reported?
  • Is “AI lab” supported by registered questions, reproducible artifacts, external criticism, and kill/pivot criteria?

Realism and validity#

  • Does each realism claim map to evidence and a tier?
  • Are historical/scientific sources used within their domain?
  • Are parameter provenance, uncertainty, sensitivity, and known limitations planned?
  • Is the fictional setting protected from implicit cultural stereotypes?
  • Are the initial population and boundary-world assumptions coherent?

Construction and institutions#

  • Can novel projects be more than reskinned recipes while remaining physically bounded?
  • Can ambiguous norms exist without bypassing executable authority?
  • Do elections cover eligibility, information, ballots, counts, disputes, succession, and legitimacy?
  • Can institutions fail, split, be ignored, or dissolve?
  • Can an office or generated rule escalate its own infrastructure permissions?

Conflict and hazards#

  • Is war local, logistical, costly, and consequential?
  • Can peaceful outcomes occur for structural reasons?
  • Are disease/fire/outside groups grounded and testable?
  • Are extinction and settlement abandonment distinguished?
  • Does the design avoid sensationalizing violence?

Architecture and operations#

  • Is the event store genuinely authoritative and replayable?
  • Are module boundaries and transactional invariants sufficient?
  • Can provider timeouts, retries, budget exhaustion, migrations, and crashes corrupt history?
  • Is a seven-day unattended run realistically operable?
  • Can renderer/UI/director accidentally become a second truth source?

Security#

  • Can resident text inject prompts or action definitions?
  • Can generated actions reach network, shell, filesystem, secrets, or arbitrary code?
  • Are authorization, idempotency, expected version, and bounds enforced server-side?
  • Are private memories/rationales separated and redacted?
  • Is public sharing/data retention addressed?

Human data and research ethics#

  • Does any page, form, issue template, API, or dataset solicit identifiable life histories or third-party nominations?
  • If human-participant research exists, is intake first-party and private, with informed consent, review, minimization, withdrawal/deletion, retention, access control, and incident handling?
  • Are vulnerable participants, relatives, bystanders, impersonation, grief, and posthumous-data risks explicitly addressed?
  • Does the project forbid telling a participant or family that an emulation is the person or that continuity has been established?

Traceability matrix#

The auditor should verify and extend this matrix.

Promise State Mechanism Evidence/UI Test/gate
Gather resources spatial stocks, lots, body, tools extraction/regrowth/activity map, lot ledger, causal event conservation + depletion / Gate 1
Build own things blueprint, materials, task DAG, skill, structure propose→compile→authorize→work→test project page/replay adversarial blueprints / Gate 3
Hold elections charter, roll, ballots, office capabilities staged election lifecycle institution/election page tie/fraud/succession / Gate 4
Create government rules, institutions, roles, legitimacy group and policy grammar charter lineage multiple forms/ablation / Gate 4
Form factions memberships, goals, identity, coordination invitations, meetings, collective actions social/group graph persistence/fission scenarios / Gate 4
Trade/share/steal lots, custody, claims, obligations transfers and detection ledger/person timeline conservation/private info / Gate 3
Learn/invent skills, techniques, teaching/prototypes practice, teaching, recombination/test knowledge lineage prerequisite/novelty tests / Gate 3+
Disease body/pathogen/contact state introduction/transmission/progression/care health/contact traces route/fade-out tests / Gate 5
Fire/disaster fuel, weather, structures, hazard state ignition/spread/damage map/causal replay zero-fuel/wind scenarios / Gate 5
War grievances, groups, supplies, participants, encounters escalation/mobilization/logistics/combat/aftermath conflict timeline/why peace/war sensitivity / Gate 5
Migration/fission membership, route, supplies, destination departure and boundary resolution map/group history no-teleport/conservation / Gate 5
Extinction living people/membership/settlement occupancy mortality/migration + terminal classifier preserved final report classification/replay / Gate 5/7
Return after 2 days event log, projections, significance read-only director Since-you-left dashboard factuality + soak / Gate 6/7
Explain why causal parents, observations, RNG, commands provenance graph Why page causal completeness sampling / all
Run within tokens activations, model profiles, usage/caps tiering/routing/degradation, cache-aware batching cost dashboard, cache hit rate budget chaos tests / Gate 2/3/7
Reproduce a history event log, manifest, response cache, state digests pinned-environment replay, decision barrier replay status, divergence report determinism suite / Gate 1–2
Survive as a population vital rates, migration, exogamy, kin graph demography plus boundary-world exchange population page, lineage view null-model comparison / Gate 1/3
Claim something emerged pre-registered definitions, founding manifest declared/randomized initial-condition cohorts, ablations emergence audit report proxy audit, seed-selection disclosure, no-charter control / Gate 3–5
Behave like distinct people values, biography, per-resident routing diversity metrics, prompt/model variation person pages, divergence stats homogeneity check / Gate 2
Preserve operational identity commitments, memories, relationships, policies, model/history manifest perturbation/recovery and transplant protocol blinded scores and failure traces R0 reliability/confound gate + Gate 2
Make a research claim registered hypothesis, conditions, seeds, exclusions, metrics controlled cohort + baselines/ablations experiment card, raw outputs, analysis independent reproduction / each claim
Protect human subjects consent record, minimized private data, retention/withdrawal state separate reviewed intake and access workflow private participant portal/audit log ethics and security review / before collection

Specific red flags#

The auditor should issue at least a High finding if any of these appear:

  • “the LLM decides what happens” without kernel transition rules;
  • “random event” without a conditional hazard definition;
  • global memory or full-state prompts for residents;
  • event replay dependent on calling a live model again;
  • currency, elections, police, war, or religion installed as assumptions despite emergence claims;
  • observer/editor intervention hidden from canonical lineage;
  • outcome-tuned “interestingness” feeding back into the world;
  • arbitrary code/shell/browser tools given to residents;
  • numerical historical claims without provenance/confidence;
  • irreversible actions without idempotency/version checks;
  • a 3D engine treated as canonical physics before Gate 6;
  • an unqualified "bit-for-bit deterministic" claim over floating-point state;
  • replay that would call a live model on a cache miss instead of failing;
  • a mechanism borrowed from a population thousands of times larger, without a validity-at-n=18 argument;
  • a collapse or cascade reported without comparison to a null demographic model;
  • a founding fixture that names a future tension, or an operational definition written after the runs;
  • a behavioral finding reported without the model's refusal rate;
  • a repository cited as upstream that turns out to be a fork, or a license claimed from a README rather than a LICENSE file.

Sign-off protocol#

  1. Reviewer writes the report without editing the plan.
  2. Plan author responds in AUDIT-RESPONSE.md, accepting, disputing with evidence, or deferring each finding.
  3. Accepted fixes update the wiki and decision log.
  4. Reviewer rechecks Blocker/High findings and updates verdict.
  5. Owner explicitly approves Gate 0 and the Phase 1 spend/deployment constraints.

“Both agents are satisfied” means no open Blocker/High finding. Medium/Low findings may remain only with owner-visible disposition and a target gate.

Part III

Evidence, integrity, and the lab

Sources, determinism, parameters, costs, initial conditions, research charter, and R0 protocol.

Research and hardeningwiki/13-sources.md
3,526 words
Inside this chapter

13 — Sources and Research Notes#

Research snapshot: 2026-08-01. Every URL below was fetched and checked during the audit pass on that date unless marked otherwise. Repository activity, prices, APIs, and licenses change; verify again at implementation.

Sources below informed architecture and mechanism selection. Inclusion is not endorsement, and is not evidence that a model is valid for December.

How to read the license column#

The licensing question in 11 — "permissive-only code reuse versus GPL-compatible project" — can now be answered from evidence. See Licensing consequences at the end of this page.

Long-horizon generative societies#

  • Agentopia repository — long-horizon multi-agent simulation; Plan→Contact→Activity→Review, self-managed memory, append-only JSONL records, checkpoint resume. License caveat: the README states MIT but the repository contains no LICENSE file, so GitHub reports it as unlicensed. Treat as legally unlicensed for reuse purposes. Also note the repository is three commits deep, all dated 2026-06-05; the headline experiment is a paper claim, not something reproducible from the four files present.
  • Agentopia paper — "Agentopia: Long-Term Life Simulation and Learning in Agent Societies," Xintao Wang et al., 2026-06-05. Source of the token figures used in 16: Table 4 reports 13,347 M input tokens, 352 M output tokens, and 567 K LLM calls, as averages across three simulated worlds (per-run input ranged 9,699–19,041 M), for 100 agents over 10 simulated years in ~186 wall-clock hours. The models used were Qwen3.5-397B with Gemini 3 Flash as a fallback for invalid outputs. Very recent; treat replication maturity cautiously.
  • Generative Agents paper — Park et al. Foundational memory-stream, reflection, planning, and believability evaluation. 25 agents.
  • Concordia repository — Apache-2.0, actively maintained. Configurable generative social simulation with a Game Master entity and modular components.
  • AI Town repository — MIT, maintained. Convex/React/PixiJS browser town; product and UI reference only.
  • MiroFish repositoryAGPL-3.0, very active. Self-described as a "swarm intelligence engine" for prediction rather than a social-simulation framework; the seed-material → parallel simulation → report-generation workflow is as described. The AGPL obligation is materially stronger than any other repository on this page and extends to network use — do not vendor code from it without a deliberate licensing decision.
  • Project Sid repositorycontains no code. It is a technical-report PDF, README, image, and video, with no license and no pushes since 2024-11-04. The PIANO architecture is described only in the paper. Useful as a claim landscape; it is not an implementation reference.
  • AI Agents Alone Are Not (Yet) Sufficient for Social Simulation — Yiming Li & Dacheng Tao. Argues role-play plausibility is not behavioral validity and proposes an environment-involved formulation. Direct support for December's kernel-owns-truth stance.
  • MemFail — "MemFail: Stress-Testing Failure Modes of LLM Memory Systems," Garg et al., 2026-05-26. Correction to earlier drafts: its taxonomy is summary failure, storage failure, retrieval failure, and reasoning failure, evaluated over five adversarial datasets. Earlier drafts of this wiki attributed to it a different list ("fabricated recall, over-compression, source confusion, self-reinforcing summaries") which is our paraphrase, not the paper's. 04 has been corrected to use the paper's actual categories.

Personal identity, agent continuity, and consciousness boundaries#

Embodied worlds and construction#

  • Mindcraft repository — MIT, active. Minecraft/Mineflayer LLM agents; providers include OpenRouter. The repository moved: the widely cited kolbytn/mindcraft now redirects here. Code execution via allow_insecure_coding must remain disabled — see ADR-005. A separate community fork, mindcraft-ce/mindcraft-ce, is more recently active but is not canonical.
  • Mineflayer repository — MIT, very active. The Minecraft bot API underlying Mindcraft.
  • Craftium repositoryLGPL-2.1-or-later for code plus CC BY-SA 3.0 for media, inherited from Luanti. Not MIT — earlier drafts implied a permissive license and were wrong. Gymnasium and PettingZoo APIs, with explicit client/server synchronization for slow agents such as LLMs. Published at ICML 2025 (paper), so it is no longer merely an early experiment — but development has slowed, with no pushes since 2026-02-17.
  • Luanti repositoryLGPL-2.1-or-later code, CC BY-SA 3.0 media with several exceptions. Formerly Minetest. Very active.
  • MineLand paper — multi-agent Minecraft simulator with limited senses and physical needs.
  • VillagerAgent paper — DAG-based task decomposition and coordination; Findings of ACL 2024.
  • APT paper — text-to-blueprint architectural planning. arXiv preprint, no venue listed.

Governance and social dilemmas#

  • GovSim repository and paperthis is the correct citation. "Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents," Piatti et al., NeurIPS 2024. MIT. Effectively a frozen research artifact (no pushes since 2025-01-19).
  • GovSimElect / AgentElectcorrection to earlier drafts, which cited this as though it were an upstream project. It is a five-star personal fork of GovSim that renames itself "AgentElect" in its own README, comparing elected-leader, fixed-leader, and leaderless governance over the GovSim fishery scenario. Cite it as a fork if the election variant is what is wanted; cite GovSim above for the underlying work.
  • Agent Ballot Box — MIT. Weak citation, retained only for completeness: zero stars, single author, no activity since 2025-06-22, and it is a GovSim-style commons framework (fisheries, pasture, pollution) in which voting is one component rather than the focus. Earlier drafts described it as "LLM voting experiments," which oversells it.
  • Melting Pot repository and Melting Pot 2.0 paper — Apache-2.0, active. The "over 50 substrates / over 256 scenarios" figures come from the repository README and full report, not the paper abstract — cite the repository for them. The README's own dilemma list is "cooperation, competition, deception, reciprocation, trust, stubbornness"; coalition-formation substrates exist but "coalitions" was our paraphrase.
  • Open Policy Agent documentation — Apache-2.0, CNCF graduated. Policy-as-code reference for infrastructure authorization; not a model of social law.

Archaeology, subsistence, and demography#

  • MoralAgentSim and paper — "Why Are We Moral? An LLM-based Agent Simulation Approach to Study Moral Evolution," Zhou et al., ACL 2026 Main (oral). The "MoRE" architecture in a prehistoric hunter-gatherer environment with hunting, gathering, sharing, communication, reproduction, and conflict. The closest published analogue to December's cognition-in-a-subsistence-world premise; assess independently.
  • AgModel at CoMSES — Isaac Ullah, v1.0.0, 2024-12-06. GPL-2.0. Forager–farmer transition with demography, environment, subsistence, labor, stores, births, and deaths. Parameter and mechanism reference.
  • Agent-Based Modeling for Archaeology — Romanowska, Wren & Crabtree, SFI Press. CC BY-SA 4.0. Ten chapters with NetLogo code covering subsistence, population, fission–fusion, commons, and games.
  • Simulating Forager Mobility — NetLogo models for residential mobility, waterhole tethering, and logistical mobility.
  • Artificial Anasazi modelboth cautions confirmed verbatim on the page: "This model is unverified. It has not yet been tested and polished as thoroughly as our other models," and "environmental variability alone can not explain the population collapse around 1350." Use as a cautionary case, never as historical ground truth.
  • Lewis et al. 2014, high mobility and demand sharingNature Communications 5:5789. Hunting, movement, sharing, reproduction, aging, and enforced cooperation in egalitarian hunter-gatherers.

Demographic and energetic anchors are consolidated in 15, which flags which values were verified in this pass and which still require URL confirmation.

Disease, ecology, and hazards#

The disease literature was researched in depth during this audit because it materially changed the design. Full treatment and design consequences are in 15 §D.

  • Bartlett 1957 and Bartlett 1960 — origin of critical community size; measles CCS 250,000–300,000.
  • Black 1966 — island study; CCS 300,000–500,000, with transmission breaks in every community below 500,000.
  • Keeling & Grenfell 1997 — shows realistic infectious-period distributions reproduce the observed CCS, giving the empirical result a mechanistic basis.
  • Black 1975, Infectious Diseases in Primitive Societies — serosurveys of isolated Amazonian groups; chronic and latent infections are endemic, acute infections die out after introduction.
  • Wolfe, Dunavan & Diamond 2007 — five-stage animal-to-human pathogen model; crowd diseases require "at least several hundred thousand people."
  • Amazonian contact epidemics, 1875–2008 — 117 epidemics across 59 societies; median mortality 18%, range <1–97%, median affected population 180. The closest empirical analogue to December's scale.
  • Faroe Islands 1846 — Panum's virgin-soil measles study; 77.5% attack rate, 2.8% CFR, lifelong immunity demonstrated over 65 years.
  • Tristan da Cunha respiratory epidemics — a ~300-person community; every epidemic ship-initiated, and a three-week passage was enough for the virus to die out en route.
  • Stochastic fade-out in branching processes — branching-process basis for early fade-out probability. It is intuition for ensemble behavior, not proof that an 18-person outbreak must fall into exactly two final-size bins.
  • Delamater et al. 2019, complexity of R₀ — R₀ is a function of social organization, not a pathogen constant; measles estimates span 3.7–203.3.
  • Caulfield et al. 2004 — undernutrition attributable fractions for child mortality, 44.8% (measles) to 60.7% (diarrhea). The quantitative basis for coupling December's subsistence and health modules.
  • Covasim and methods paper — MIT. The repository has moved to starsimhub/covasim. Adapt architecture and testing discipline, never COVID parameters.
  • OpenABM-Covid19 paper — individual contact networks and extensive test categories.
  • ForeFire documentationGPL-3.0. Wildfire fuel/landscape/spread concepts; far more detailed than v1 needs.

Agent-based-model engineering and validation#

  • Mesa documentation — Apache-2.0. Current stable is 3.5.1 (2026-03-15), with a 4.0 line in alpha. Correction to earlier drafts: discrete-event and hybrid step+event scheduling are now stable, not experimentalmodel.schedule_event(), model.schedule_recurring(), run_for, run_until. The mesa.experimental.devs simulators (Simulator, ABMSimulator, DEVSimulator) are deprecated since 3.5.0 and removed in the 4.0 line, so any design targeting them is already stale. Note also that mesa.space is maintenance-only and SolaraViz carries breaking-change risk across minor releases — both argue for the narrow SimulationRuntime interface in 02.
  • ODD protocol resource and ODD 2020 update — Grimm et al., JASSS 23(2):7, DOI 10.18564/jasss.4259.
  • USGS overview of the ODD second update — indexed at usgs.gov/publications/odd-protocol-describing-agent-based-and-other-simulation-models-a-second-update but fetch-blocked during this audit; the pubs.usgs.gov record 70209554 also returns 403. Use the JASSS version above as the citable source.
  • Azure event-sourcing pattern — general event-store, replay, and projection tradeoffs. Insufficient on its own: it does not cover the commit-ordering hazard that 14 §D7 addresses.

Determinism, state integrity, and replay#

Added by this audit pass; these underpin 14.

Open-source strategy and economy references#

  • 0 A.D.GPL code, CC BY-SA 3.0 art. Latest is Release 28 "Boiorix" (2026-02-18); the project dropped "Alpha" branding at R28. Reference only; rejected as an authoritative base.
  • Widelands and unit documentationGPL-2.0, actively maintained (v1.3.1, 2026-02-22). Worker/ware/building economy and moddable Lua unit definitions.
  • Unknown Horizonsdormant, not maintained. Correction to earlier drafts: the last release is 2019-dev (2019-01-13); the classic game is Python/FIFE with no LICENSE file in the repository, and the Godot rewrite (godot-port, GPL-2.0) has no playable content and no release. "Python/Godot" conflates two separate codebases.
  • FreecivGPL-2.0, very active (3.2.5, 2026-07-10). Ruleset and client/server architecture reference only.

Model routing, telemetry, and provider facts#

Licensing consequences#

Evidence now supports a concrete answer to the open licensing question in 11:

If December is… Then these are usable And these are not
Permissive-first project Mesa (Apache-2.0), Melting Pot (Apache-2.0), Concordia (Apache-2.0), GovSim (MIT), Mineflayer/Mindcraft (MIT), Covasim (MIT), OPA (Apache-2.0), permitted open-core portions after file-level review LGPL/GPL/AGPL code requires deliberate integration and distribution/network-use analysis; it is not categorically unusable
Reciprocal-license-compatible project May incorporate more LGPL/GPL components while satisfying their terms AGPL and mixed-media licenses still require specific review, especially for a network service

Three practical notes:

  1. The 3D adapter decision at Gate 6 is also a licensing decision. Craftium and Luanti are LGPL-2.1-or-later with CC BY-SA media. Dynamic linking and process separation keep LGPL obligations manageable, and December's adapter contract in 06 already implies process separation — but this must be a deliberate, recorded choice rather than a discovery made at integration time.
  2. Reading a GPL model for mechanism inspiration is not reuse. AgModel and the archaeology ABMs are parameter and mechanism references; that use is unaffected by their license. Only copied code creates an obligation. Record which is which in the parameter registry.
  3. Do not treat this table as legal advice. MiroFish’s AGPL is especially relevant to network deployment, but LGPL linking, GPL distribution, assets, open-core carve-outs, and combined works all need project-specific review.

Source-quality rules for implementation#

  1. Prefer primary papers, official documentation, official repositories, and maintained model registries.
  2. Record version, commit, date, and license for any code or parameter reuse.
  3. Do not transfer a parameter outside its population, environment, or domain without stating the justification.
  4. Triangulate important numerical values across multiple sources, or carry a broad uncertainty range.
  5. Separate mechanism inspiration from empirical calibration.
  6. Mark unverified models and very recent papers explicitly.
  7. Preserve a bibliography entry beside every parameter family in the registry.
  8. Verify that a cited repository is the upstream project and not a fork, and that a claimed license is backed by a LICENSE file rather than a README sentence. This audit found errors of both kinds.
Research and hardeningwiki/14-determinism-replay-and-state-integrity.md
3,320 words
Inside this chapter

14 — Determinism, Replay, and State Integrity#

Status: added by the 2026-08-01 audit pass. This document replaces the informal determinism language in 00, 06, and ADR-002 with a specification that survives contact with real hardware.

Why this document exists#

Earlier drafts claimed replays are "bit-for-bit identical in the non-LLM kernel." That claim, unqualified, is false for any Python simulation that uses floating-point arithmetic, and it is the load-bearing claim under event sourcing, shadow worlds, the causal UI, and the audit protocol. If replay is not actually reproducible, every downstream integrity promise is decoration.

This document states what is achievable, at what cost, and what the project commits to.

D1. Three reproducibility guarantees#

Earlier versions used “replay” for three different operations. They are now separate:

  1. Recorded-history reconstruction. Apply the already-recorded canonical events to an empty projection or compatible snapshot. This must be portable across supported machines because events contain the authoritative outcomes and use canonical encodings. No weather, RNG, physics, or model decision is recomputed.
  2. Kernel re-execution. Starting from the same initial state, commands, RNG keys, code, and config, recompute the events. Exact checkpoint equality is required inside the pinned numeric environment. Portable equality is a Phase 1 goal if canonical state uses fixed units and deterministic transition implementations.
  3. Fresh counterfactual simulation. Re-run one or more cognition decisions against a live model or change an intervention. This is a new branch, never a replay, and is expected to diverge.

Claim. Recorded history reconstructs exactly from its event artifacts. Kernel re-execution reproduces checkpoint hashes within the declared environment. Fresh model sampling produces a separately identified branch.

This distinction prevents floating-point caveats from weakening ordinary event-store recovery while keeping re-simulation claims honest.

Why portable bit-identity is not available#

IEEE 754 requires correctly rounded results for + - * / sqrt fma, so a fixed sequence of basic operations is reproducible across conforming hardware. Everything that breaks reproducibility is about which sequence actually executes:

Divergence source Effect Applies to us?
FMA contraction (a*b+c as one rounding vs two) Different low bits; x86 and ARM64 toolchains default differently Yes, if the kernel does float arithmetic in compiled extensions
SIMD reduction order (SSE/AVX2/AVX-512/NEON lane counts) Reassociated sums differ Yes, via NumPy's runtime CPU dispatch
libm transcendentals (exp, log, pow, sin) glibc explicitly disclaims correct rounding; results differ by version, OS, and architecture Yes — this alone falsifies portable bit-identity
BLAS backend (OpenBLAS vs MKL vs Accelerate) and its thread count Different results for any linear algebra Only if the kernel uses @ or np.linalg
-ffast-math / FTZ / DAZ flags, including flags set by other loaded libraries Reassociation and denormal flushing Possible, via third-party wheels
x87 80-bit intermediates Extra precision on 32-bit x86 No — we require SSE2/NEON code paths

NumPy makes no cross-platform or cross-version guarantee. np.sum uses pairwise summation "when the summation is along the fast axis in memory," so its result depends on array layout, axis, and strides, and on the SIMD width selected at runtime. The NumPy random compatibility policy is explicit that identical streams are promised only for "the same build of numpy, in the same environment, on the same machine."

Correctly-rounded math libraries are maturing — the CORE-MATH project has been partially upstreamed into glibc 2.42/2.43, and LLVM libc implements C23 math correctly rounded — but relying on whatever libm the host ships remains unsound. Portable determinism would require bundling a correctly-rounded math layer or making the kernel integer-only.

Numeric policy for kernel re-execution#

Phase 1 must answer one question: can the canonical kernel state be integer-valued?

Conserved and discrete canonical quantities—mass, volume, energy reserve, time, inventory counts, cell coordinates, and currency if it ever exists—should use scaled integers with explicit rounding. Scientific calculations may use floats inside a transition, but only their quantized declared output enters canonical state. Phase 1 tests boundary cases across supported environments before claiming portable kernel re-execution.

This is the recommended target. It costs discipline, not performance. Gate 1 records the decision explicitly:

  • Option A (preferred): fixed-unit canonical state, with deterministic quantization at transition boundaries.
  • Option B (fallback): float canonical state and pinned-environment re-execution only.

Do not defer this decision past Gate 1. Retrofitting fixed-point arithmetic into an existing ecology and crop model is a rewrite.

D2. Kernel execution constraints#

Whichever option is chosen, the kernel obeys these rules. They are testable and belong in CI.

  1. The kernel is a synchronous, single-threaded pure function of (prior state, ordered input events). Concurrency lives strictly outside it. No asyncio inside a transition rule, no thread pools, no executor callbacks.
  2. Run on the default GIL build. Free-threaded CPython is officially supported as of 3.14, but true parallelism reintroduces allocation-order and interleaving nondeterminism. If free-threading is ever adopted, the kernel stays confined to one thread with no shared mutable state.
  3. PYTHONHASHSEED is pinned in the world manifest and asserted at startup. str/bytes hashing is salted per process by default.
  4. Never iterate an unordered collection. dict preserves insertion order (guaranteed since 3.7) and is permitted when insertion order is itself deterministic. set and frozenset iteration is forbidden in any code path that writes state; use sorted(collection, key=<stable canonical key>). A lint rule enforces this.
  5. No id()-derived ordering and no default object.__hash__ for any entity that participates in state. Entities sort by their stable string ID.
  6. No finalizers. __del__, weakref callbacks, and resurrection are banned in kernel modules; GC timing is not reproducible.
  7. No wall-clock reads inside the kernel. Simulated time is a state variable. time.time(), datetime.now(), and loop.time() are banned in kernel modules and blocked by lint.
  8. No platform libm in state-writing code under Option B. If a transition rule needs exp or pow (crop growth curves, decay, hazard rates), it uses a project-owned implementation: a table-plus-polynomial approximation with declared precision, versioned as a parameter. This is a small amount of code and it removes the single largest portability hazard.
  9. Pin SIMD dispatch and BLAS under Option B: set NPY_DISABLE_CPU_FEATURES to a documented baseline, pin the NumPy wheel in the lockfile, and keep linear algebra out of the kernel entirely.
  10. Canonical serialization for hashing: fixed field order, integers where possible, and for any float a fixed-width byte encoding — never repr().

D3. Random number generation#

Use counter-based, stateless, key-derived streams. A counter-based generator is a pure function output = f(key, counter), so it needs no sequential state in snapshots, can be evaluated from any point, and gives independent streams by construction.

stream_key = SHA256(root_seed ‖ branch_id ‖ subsystem ‖ entity_id ‖ purpose)[:16]
draw       = Philox(key=stream_key, counter=logical_draw_index)

Rules:

  • Never seed = hash(entity_id). Python's hash() is salted, so it is not even stable across processes; truncated seed spaces invite birthday collisions between entities; and related seeds produce correlated streams in some generators.
  • Use the raw BitGenerator bit stream as the stable layer. NumPy's NEP 19 policy permits breaking distribution streams (normal, poisson, multivariate_normal) in minor releases for performance or correctness. Integer bit output from PCG64/Philox is stable in practice; the float distribution layer is where version risk concentrates. December implements its own inverse-CDF transforms over raw bits for any distribution whose stream must survive a dependency bump, and records the distribution implementation version on every stochastic event.
  • One stream per (subsystem, entity, purpose). This delivers the property 06 already demands: adding a cosmetic draw cannot perturb crop yields, because cosmetic draws live in a different stream with a different key.
  • Record distribution_version, parameters, counter, and sampled result on every stochastic event, so a draw can be re-derived and audited without re-running the subsystem.
  • Pin the NumPy version in the world manifest regardless.

D4. Laundering LLM latency into deterministic order#

This is the mechanism 07 asserts ("no resident gains advantage from model-provider response latency") without specifying. The pattern is standard in deterministic simulation testing and lockstep game networking: the kernel never observes arrival order; arrival order is laundered into a logical-time decision that is recorded as an event.

The decision-barrier protocol#

  1. Emit. Kernel tick at simulated time T ends by emitting DecisionRequested events. Each carries a deterministic decision_id = H(world, branch, T, actor_id, trigger_kind, sequence_within_tick). The kernel's work for tick T is now finished; it does not wait.
  2. Dispatch. An orchestrator outside the kernel resolves each request: check the content-addressed response cache, otherwise call the provider. Latency, retries, provider failover, and concurrency all happen here and are invisible to the kernel.
  3. Resolve. Each response is appended as a DecisionResolved event bound to a logical barrier — either a fixed offset T+k (turn buffering, as in classic lockstep RTS) or "the first tick after all requests issued at T have resolved or timed out." The choice is a config parameter; both are deterministic.
  4. Ingest. At the barrier tick the kernel consumes all resolved decisions sorted by decision_id, never by arrival time, and validates each through the normal command pipeline.
  5. Resolve live failure explicitly. A wall-clock deadline and retry budget may decide whether a response is accepted in the live service. The first durable resolution wins. Low-significance requests may record a resident-specific routine fallback; unresolved high-significance requests pause the world at that barrier. Recorded-history reconstruction consumes this resolution and never reruns the race.

Consequence: network arrival order never breaks ties between already-resolved resident actions. Provider failure may affect wall-clock pace and whether the live world pauses, and that operational history is visible rather than laundered away.

Decision-level idempotency#

Command-level idempotency keys (already in 06) prevent a duplicate delivery of the same command from applying twice. They do not cover the more likely failure: a provider call times out, the orchestrator retries, and the model returns a different command the second time. Both are validly signed, neither is a duplicate, and the world silently forks from the replay.

Rule: the decision_id is the idempotency unit, not the command. The first response to be durably recorded for a given decision_id wins permanently. Late-arriving responses for a resolved decision_id are discarded and counted in a metric; they never reach the kernel. This must be enforced by a unique constraint on decision_id in the resolution table, not by application logic.

D5. The model-response cache is canonical, not an optimization#

No provider offers reproducible sampling. As of August 2026: OpenAI's seed is documented best-effort with no guarantee; Gemini's seed is likewise best-effort; Anthropic has no seed parameter at all and its newest models reject temperature/top_p entirely; DeepSeek and MiniMax expose no reproducibility mechanism. Temperature 0 is not deterministic on any of them.

Therefore:

  • The content-addressed prompt/response cache is part of the canonical world artifact, backed up and versioned with the event log. It is not a disposable performance cache.
  • The cache key is the full request as sent: prompt bytes, model ID and provider-reported version/fingerprint, tool/schema definitions, and every sampling parameter.
  • Replay hard-fails on a cache miss. It must never silently call a live model — that would fabricate a divergent history under the banner of reproduction. A replay that encounters a miss aborts and reports the missing decision_id. This is a required test.
  • Re-running with fresh model responses is not a replay. It is a new branch with its own branch_id, and the UI must label it as such.
  • Prompt/response bodies are stored encrypted at rest where feasible and are subject to the retention policy in 08.

D6. State hashing that does not cost O(state) per event#

06 requires pre_state_hash and post_state_hash on every event. Implemented naively — serialize the world, hash it — this is O(total state) per event and makes the event pipeline quadratic in world size. With 10⁴–10⁶ entities and millions of events it is not merely slow, it is the dominant cost of the entire simulation.

Phase 1 minimum: hash-chain event envelopes, keep expected aggregate versions, compute hashes for changed aggregates, and compute a canonical full-state hash at snapshots. This is sufficient for reconstruction, corruption detection, and divergence localization at December’s starting scale.

Only if measurement shows snapshot hashing or aggregate comparison is too slow should December add incremental homomorphic hashing. An LtHash-style construction would maintain a running digest under addition and subtraction:

H' = H − h(entity_id ‖ canonical_bytes_before) + h(entity_id ‖ canonical_bytes_after)

Properties that matter here: update cost is O(changed entities) rather than O(state), the digest is order-independent (so it cannot accidentally encode iteration order), and it parallelizes trivially. This is the same primitive Solana adopted to hash its entire account set every block, which is direct evidence it holds up at far larger scale than December needs.

Supplement it with:

  • Periodic full recomputation — at every snapshot, and at a configurable event interval — to catch drift and implementation bugs. A homomorphic digest that has silently desynchronized from reality is worse than no digest.
  • A Merkle tree over entity partitions, rebuilt at snapshots only. A flat homomorphic hash tells you two states differ; it cannot tell you where. When a nightly replay check fails, the Merkle tree localizes the diverging subtree in O(log n) instead of forcing a manual diff of the whole world. Divergence localization is an operational necessity, not a nicety.

Do not use verkle trees. They solve proof-size problems December does not have, and Ethereum itself has moved toward a unified binary hash tree (EIP-7864) for simplicity and post-quantum reasons.

Until that escalation, event envelopes carry the affected aggregate’s pre/post version and hash plus the previous event hash. Snapshot records carry the global state hash and algorithm version. Do not make lattice hashing a Gate 1 dependency without evidence.

D7. Reading the event log in commit order#

The sequence: 123 field in the event envelope must not be consumed with a naive high-water mark. This is the single most common way an event-sourced system on PostgreSQL loses data, and the current architecture doc walks straight into it.

The failure: bigserial values are allocated at insert time, but transactions commit in a different order. A projection reading WHERE sequence > :last_seen can observe sequence 300 before sequence 250 has committed, advance its watermark past 300, and then never see 250. Sequences are also non-transactional, so rollbacks leave permanent gaps that are indistinguishable from not-yet-committed rows. The projection is now permanently, silently wrong — and because it is silent, the "projection rebuild equals live projection" invariant in 06 is the only thing that would ever catch it.

December’s Phase 1 solution is deliberately small:

Enforced serialized writer. The canonical kernel is logically single-threaded, so one fenced leader appends events. The sequence/stream version is assigned and events are committed inside that serialized transaction. Projections may consume only committed batches and checkpoint the batch/stream version, not an independently allocated bigserial watermark.

If December later introduces multiple event writers, it must adopt and test a real commit-order protocol—such as transactional batches with a committed batch table, database commit timestamps/LSN where appropriate, or a verified snapshot-fencing design. xid8 ordering is not treated as commit ordering by itself.

The prior audit proposed the following snapshot-fencing query as defence in depth:

SELECT * FROM events
WHERE txid < pg_snapshot_xmin(pg_current_snapshot())
  AND (txid, sequence) > (:last_txid, :last_sequence)
ORDER BY txid, sequence;

It is retained as a research note, not the Phase 1 design. Transaction IDs are not commit sequence numbers, wraparound and visibility semantics require care, and the serialized writer already removes the race.

Additional operational rules:

  • LISTEN/NOTIFY is a doorbell, not a delivery mechanism. It is not durable, it is lost on disconnect, and NOTIFY at commit takes a global exclusive lock that serializes all notifying commits — a measured ceiling in the low thousands of commits per second. Use it only to wake a poll loop that is already correct without it.
  • Isolation: READ COMMITTED plus the single-writer lock and optimistic (stream_id, expected_version) concurrency. SERIALIZABLE buys nothing on a single-writer append path and adds predicate-lock overhead and retryable failures.
  • Partition the event table by simulated time once past tens of millions of rows; keep the append path's index set minimal; keep partition count in the low hundreds.
  • Expect roughly 5k events/s from a single writer on moderate hardware. See D8 — that is far above December's requirement, so throughput is not the binding constraint. Storage growth is.

D8. Event volume and storage budget#

No prior draft estimated how much history December produces, while committing to "keep canonical history indefinitely." The two must be reconciled before Phase 1 chooses a grid resolution, because grid resolution is the dominant driver of event volume.

Order-of-magnitude estimate, to be replaced by measurements at Gate 1:

Source Events per simulated day Note
Resident activity, movement, transfers, speech 300–1,500 ~25–80 per full resident
Projects and construction 50–300 Bursty; near zero between projects
Health, physiology, demography 50–200 Daily boundary settlements
World processes (weather, hydrology, crops, vegetation, prey) 100 – 40,000 Entirely determined by event granularity policy
Total ~500 – 42,000 Two orders of magnitude of uncertainty

That spread is a design decision, not an unknown. A 5 km × 5 km valley at 50 m resolution is 10,000 cells; at 25 m it is 40,000. Emitting one event per cell per day for vegetation and soil moisture produces 3.6–14.6 million events per simulated year from vegetation alone — which at ~1 KB per row is 4–15 GB per simulated year, before residents do anything.

Policy:

  1. World processes emit batched events with array payloads — one world.daily_advanced event per subsystem per day carrying a compact typed array of per-cell deltas, not one event per cell.
  2. Emit only on material change. A cell whose state is unchanged within declared tolerance emits nothing.
  3. Per-cell granularity is reserved for cells with residents, claims, structures, crops, or active hazards — typically tens of cells, not thousands.
  4. Target: ≤ 5,000 events per simulated day, verified at Gate 1. Under this policy, at 72 simulated days per real day, the log grows roughly 100–400 MB per real day, or 40–150 GB per real year. That is affordable on local disk but it is not free, and it makes the backup and restore-drill requirements in 08 load-bearing rather than ceremonial.
  5. The prompt/response cache adds its own volume — roughly 1–5 MB per simulated day after deduplicating stable prefixes. It is canonical (D5), so it is backed up too.
  6. Cold-tier archival is permitted; deletion is not. Events older than a configured horizon may move to compressed object storage with their hashes retained online, so the chain remains verifiable without the bodies being hot.

D9. Required tests#

These belong in CI from Phase 1, not in a later hardening phase.

Test Asserts Phase
Replay-from-empty golden histories State hashes match at every checkpoint 1
Replay-from-snapshot Snapshot + subsequent events equals full replay 1
Cache-miss abort Replay with a deliberately evicted cache entry fails loudly and names the decision_id 2
Latency-shuffle Same manifest, randomized artificial provider delays and completion order → identical history 2
Decision-retry divergence A timed-out call whose retry returns a different command → first recorded resolution wins; no fork 2
Cross-platform divergence Same log replayed on x86-64 and ARM64; report first diverging event, or pass under Option A 1
Unordered-iteration lint No set/frozenset iteration, no id() ordering, no wall-clock call in kernel modules 1
Hash-seed sensitivity Run under three PYTHONHASHSEED values → identical history 1
Incremental-vs-full digest Homomorphic digest equals full recomputation at every snapshot 1
Sequence-gap injection Deliberately commit events out of sequence order; every projection still sees all of them 1
Cosmetic-draw isolation Adding draws to a cosmetic stream does not perturb any physical outcome 1
Dependency-bump replay Bumping NumPy re-runs golden histories; any stream change is caught, not silently absorbed ongoing

A failure in any of these is a Blocker for the gate it belongs to. Determinism is not a property that can be added later; it is a property that is either maintained continuously or lost permanently.

Research and hardeningwiki/15-parameter-registry.md
8,435 words
Inside this chapter

15 — First-Cut Parameter Registry#

Status: provisional research notebook. Added by audit pass 1 and downgraded by pass 2 pending row-level reproducible provenance and transfer review.

Every prior draft said parameters would be sourced "during Phase 1." That deferral hid the fact that several design commitments were already numerically impossible. This document supplies first-cut anchors so the plan can be checked against reality before code is written, and so Phase 1 calibration starts from a bibliography rather than a blank page.

How to read this#

Each value currently carries an audit evidence class, not implementation approval:

Class Meaning
V Auditor-reported verification — not reproducible provenance without a source ID and exact locator on the row
D Derived — computed by the audit from verified inputs; arithmetic, not measurement
S Secondhand — a paywalled primary quoted verbatim by a fetched peer-reviewed source
C Contested — the literature genuinely disagrees; must be a range or scenario axis, never a point value
N Not verified — could not be confirmed; do not use in code without checking

No value here is calibrated or approved for December. These are research leads: they may bound plausibility or reveal nonsense, but most come from populations and environments unlike the fictional Founding Valley. Before a row enters code it must gain source_id, exact table/page/section locator, source population, transformation formula, intended subsystem, transfer justification, uncertainty representation, and reviewer status. A bold number plus “V” is not enough.

Interpretive prose in this notebook is hypothesis, not design law. Phase 1 admits only the smallest parameter subset needed for its registered patterns; genetics, warfare history, and other later-domain material remain out of the kernel until their phase and research question justify them.

Read the design notes, not just the tables. Several numbers below carry consequences that change how a subsystem must be built.


A. Human energetics#

Parameter Central Range Unit Context Class
PAL, sedentary/light 1.55 1.40–1.69 ×BMR FAO/WHO/UNU 2004 Table 5.3 V
PAL, moderately active 1.85 1.70–1.99 ×BMR same V
PAL, vigorous 2.20 2.00–2.40 ×BMR same; >2.40 unsustainable long-term V
PAL, non-mechanized agriculture — FAO category 2.25 2.00–2.40 ×BMR FAO worked example V
PAL, farmers — actually measured M 1.90 / F 1.74 M 1.36–2.40 / F 1.47–2.36 ×BMR Doubly-labelled-water review, 26 studies V
PAL, Hadza foragers M 2.26 / F 1.78 ±0.48 / ±0.30 ×BMR DLW, n=30 V
PAL, survival minimum 1.27 ×BMR Totally inactive dependent person V
BMR, Schofield M 18–30 15.057W + 692.2 see 153 kcal/day W in kg; FAO 2004 retains Schofield V
BMR, Schofield F 18–30 14.818W + 486.6 see 119 kcal/day same V
BMR, Henry M 18–30 16.0W + 545 kcal/day UK SACN 2011 adopted Henry V
BMR, Henry F 18–30 13.1W + 558 kcal/day same V
Henry vs Schofield −3 to −4 % BMR Henry lower; 79% vs 69% within ±10% of measured V
TEE, man 60–70 kg, subsistence agriculture 3,200 2,900–3,900 kcal/day DLW measured; upper bound = peak harvest season V
TEE, woman 50–60 kg, subsistence agriculture 2,400 2,100–2,800 kcal/day DLW measured V
Walking cost, level, gross 0.81 0.72–0.91 kcal·kg⁻¹·km⁻¹ 3.4 J/kg/m; meta-analysis of 13 studies V
Walking cost, level, net 0.57 0.38–0.67 kcal·kg⁻¹·km⁻¹ 2.4 J/kg/m V
Load carriage M = 1.5W + 2.0(W+L)(L/W)² + η(W+L)(1.5V² + 0.35VG) watts Pandolf 1977; G in percent, V in m/s V
Terrain factor η 1.0 blacktop / 1.1 dirt / 1.2 light brush / 1.5 heavy brush / 2.1 sand ±0.1–0.2 USARIEM's own field data gave 1.2 for "1.1" surfaces V
Minnesota: semi-starvation 24 weeks @ ~1,570 kcal/day 36 men started, 32 analysed V
Minnesota: weight loss −24 −25 target % body mass 69.4 → 52.6 kg; BMI 21.9 → 16.4 V
Minnesota: strength decline −28 to −37 % Grip dynamometer V
Minnesota: work capacity decline −72 to −94 % Harvard fitness test V
Minnesota: VO₂max decline >−40 % At 25% weight loss V
Minnesota: BMR decline −38 15–25 pts adaptive % 1,608 → 994 kcal/day V
No performance loss below 10 % body-mass loss Taylor 1957 V
Lethal BMI, men / women 13 / 11 survival below 10 documented kg/m² Henry 1990; ~21 fatal cases — modal, not a wall S, C
Body-mass loss at death 30 25–35 % initial Hunger-strike series S
Starvation survival, water available 62 46–73 days 10 IRA hunger strikers, healthy adult males V
Starvation survival, broader series 62.5 11–115 days 20 political strikers since 1920 V
Survival without water, 32 °C shade, resting 3 2–5 days Adolph's table S
Survival without water, 21–27 °C 10 9–12 days same S

Design note A-1 — starvation is a behavioral parameter, not only a physiological one#

The Minnesota figures are the most useful thing in this section, and the useful part is not the weight loss. Work capacity fell 72% while body mass fell only 24%. An undernourished resident is not a slightly slower worker — they are a person whose capacity to do anything demanding has collapsed nonlinearly, well before they look like they are dying.

The behavioral record is equally specific and equally implementable: 97% reported tiring easily; two-thirds were downhearted with concentration difficulty; there was "no sign of a drive for activity"; food preoccupation dominated waking thought; sexual interest and sociability withdrew. Notably, measured intellectual ability, memory, and logic did not decline even though subjects believed they had — a distinction worth preserving, because it means a starving resident should still reason competently while wanting almost nothing except food.

This belongs in the appraisal prompt and in the behavioral test suite (09). It is one of the few places where a well-documented human response can directly discipline LLM behavior rather than leaving "acts hungry" to the model's imagination.

The cited performance threshold must not be turned into “hunger is free below 10% weight loss.” Hunger, attention, mood, and food-seeking can change before measurable strength/work-capacity loss, and the study population does not justify a universal threshold. Model immediate hunger/appraisal pressure separately from a delayed nonlinear impairment curve, and sweep the curve under explicit uncertainty.


B. Subsistence production#

Parameter Central Range Unit Context Class
Wheat yield ratio 3.66 2.16–5.35 (p10–p90) seed:seed 8,403 English manor-years, 1211–1491 V, D
Barley yield ratio 3.60 2.12–5.19 seed:seed 7,439 manor-years V, D
Oats yield ratio 2.56 1.41–3.75 seed:seed 9,479 manor-years V, D
Rye yield ratio 3.92 2.03–6.00 seed:seed 1,247 manor-years V, D
Harvest retained as seed, wheat 27 19 (good yr) – 46 (bad yr) % = 1/ratio D
Harvest retained as seed, oats 39 27–71 % D
Contemporary break-even ratio 3.0 seed:seed Walter of Henley's medieval accounting target V
Wheat yield, absolute 554 327–810 kg/ha gross ratio × sowing rate D
Emmer, hand-hoed, no manure 2.08 0.78–3.11 t/ha Butser Ancient Farm, 15 consecutive seasons V
Emmer, swidden, years 1→4 1.2 → 1.1 → 0.7 → 0 t/ha Butser slash-and-burn, mattock hoe V
Millet, low-input smallholder 500–1,500 150–3,000 kg/ha Modern African analogue V
Yield CV, individual field-year 0.36 0.36–0.41 Across four cereals D
Yield CV, regional annual mean 0.14 0.11–0.15 National annual means D
Complete crop failure frequency 7 % of seasons Butser: 1 in 15, killed by frost V
P(wheat ratio < 2.0) 7.0 % of field-years Half or more of harvest must be re-sown D
Lag-1 autocorrelation of yields 0.21 (wheat) 0.21–0.66 r Bad years cluster, spring grains especially D
Cross-crop correlation 0.41 0.30–0.59 r Diversification genuinely reduces risk D
Total swidden labor 175 64–330 person-days/ha 17 SE Asian cases, core operations V
Total swidden labor 1,050 384–1,980 person-hours/ha at 6 h/person-day D
— slashing and felling 41 9–75 person-days/ha 18% of total V
— burning and ground prep 12 1–47 person-days/ha 7% V
— planting 23 8–55 person-days/ha 14% V
weeding 49 0–153 person-days/ha 16% under long fallow → 38% under short V
— reaping, threshing, winnowing 50 23–73 person-days/ha 27% V
Labor productivity 6.6 3.8–25.1 kg grain/person-day V
Saddle-quern grinding 0.69 0.57–0.90 kg grain/hour Coarse meal reaches 2.1; fine meal 0.63 V
Grinding labor 1.45 1.1–1.75 hours per kg Inverse of above D
Rotary quern 3.0 kg grain/hour 4.3× the saddle quern V
Grinding energy cost 206 kcal/hour (PAR 3.5) 30 women, indirect calorimetry V
Daily grinding time 3–5 1.7–8 hours/day per woman Ethnographic, 5–10 person household V
Storage loss, cereals, 9-month season 3 1–6 % Millet 1%, sorghum 2–4%, wheat 3–5% V
Storage loss, insect-infested 11 9.7–13.3 % Maize with Larger Grain Borer V
Storage loss, first 3 months ~0 % Loss is convex in time, not linear V
Whole post-harvest chain loss 13 9–18 % field → consumption Storage proper is only 1–5 points of this V
Foraging return, men (incl. search) 1,339 1,018–1,619 kcal/hour Aché, 611 man-foraging-days V
Foraging return, women (incl. search) 1,221 302–2,804 kcal/hour Aché, 61 woman-days V
Between-forager spread 446–2,124 kcal/hour 25 Aché men — a 4.8× individual spread V
Honey, on-encounter >20,000 kcal/hour Aché V
Horticulture vs foraging 1.5–2× multiplier Return rate advantage V
Large-game hunting success 0.03 0.027–0.034 P per hunter-day Hadza, large game only; failure >97% V
Hunting success, all prey incl. small game 0.23 (!Kung) – 0.65 (Aché) 0.23–0.76 P per hunter-day Not the same quantity as above S
Hunting daily-return CV 8.1 SD/mean Hadza: mean 4.89 kg, SD 39.7, n=2,072 D
Meat package size 12.8 SD 14.0 kg Hiwi; CV ≈ 1.1 V
Gathered package size 4.3 SD 4.1 kg Aché; CV ≈ 0.95 — much tighter V
Shellfish collecting 1,492 ±173 SE kcal/hour incl. search Meriam reef flat V
Spearfishing 292 ±135 SE kcal/hour incl. search Meriam; 8.6× search penalty V
Water, drinking + water in food 2.75 2.5–3 L/person/day Sphere survival standard V
Water, total domestic 15 7.5–15 L/person/day Sphere minimum standard V
Firewood, cooking only 1.0 0.85–1.25 kg/person/day Three-stone fire, warm climate V
Firewood, cooking + indoor warming 1.8 0.5–4.0 kg/person/day Temperate winter V
Firewood, cold climate pre-industrial 4–10 up to 10 kg/person/day Northern Europe S

Design note B-1 — the seed-grain ratio is the most consequential number in the model#

At a 3.66:1 return, 27% of every harvest must be withheld from hungry people to plant next year, and in roughly one field-year in fourteen that figure exceeds 50%. Medieval accountants used a 3:1 return as the break-even line, which means a settlement is routinely operating within a factor of ~1.2 of not being able to re-sow.

This can make a famine cascade materially real rather than decorative. Eating seed grain is a possible immediate-survival tradeoff with delayed cost. It may become a private, household, or institutional issue; calling it the settlement’s “first governance dispute” would script the outcome that the emergence protocol is meant to observe. The ratio remains provisional until crop, sowing system, climate, and storage assumptions match the selected scenario.

The distribution matters as much as the mean, and it is well characterized: manor-year ratios are close to lognormal (wheat μ=1.232, σ=0.370), bad years cluster (lag-1 autocorrelation up to 0.66 for spring grains, with three-year runs of bad harvests occurring repeatedly in the medieval record), and cross-crop correlations of 0.30–0.59 mean planting two crops genuinely reduces risk rather than merely appearing to. That gives residents a real, learnable, non-obvious strategy.

Critically: the variance that matters is field-level (CV 0.36), not regional (CV 0.14). A household lives on its own field.

Design note B-2 — hunting variance can create strong risk-pooling pressure#

The Hadza coefficient of variation on daily hunting returns is 8.1. A hunter fails on more than 97% of days when large game is the target; the settlement eats meat because someone succeeds, not because anyone reliably does.

A model that gives hunters only their mean return per outing erases a major risk-pooling pressure. High variance can support sharing, reciprocity, prestige, and obligations, but it is not “the reason society exists,” and one population’s large-game rate is not a universal December parameter. Model an appropriate distribution rather than only a mean after prey set, technology, habitat, and comparison population are chosen.

Note also the prey-breadth decision, which moves success probability by an order of magnitude: 0.03/day for large game only versus 0.23–0.65/day when small game is included. Decide December's prey set first, then pick the matching rate. Do not average across these populations — they are answering different questions.

The cited datasets suggest different return distributions for gathered food and some hunting strategies. Treat the resulting sharing/obligation hypothesis as something the simulation can test, not as a universal social law.

Design note B-3 — grinding is a hidden labor sink comparable to farming#

Saddle-quern grinding can consume several household labor-hours per day under grain-heavy diets. The comparative record often shows gendered allocation, but December must not hard-code that allocation into a fictional society. Model the task burden, skill, fatigue, household negotiation, and norms; let labor allocation arise from declared initial culture and resident choices.

If December models fields and harvests but not processing, it will silently hand the settlement several free person-hours per day and misrepresent who is doing the work. Grain must be processed before it is food.

Design note B-4 — weeding is where the Boserup mechanism lives#

Total labor per hectare stays roughly flat as fallow shortens (~175 person-days), but its composition inverts: clearing falls and weeding rises from 16% to 38% of all labor, while yields decline. Labor productivity collapses even though labor input does not rise.

This is the intensification trap, empirically confirmed, and it is exactly the kind of slow structural pressure that makes a settlement's success into its later problem. It should fall out of the fallow-length state variable rather than being scripted.


C. Demography#

Parameter Central Range Unit Context Class
Life expectancy at birth 31 21–37 years Hunter-gatherers, 5 populations V
e0, forager-horticulturalists 33 21–42 years V
e0, pre-industrial Sweden 1751–59 34 years Inside the forager range V
Modal adult age at death 72 68–78 years Conditional on reaching adulthood V
Infant mortality, first year 23 21–27 % of births V, D
Survival to age 15 (l15) 0.57 0.44–0.73 proportion Foragers; forager-horticulturalists 0.64 V
Survival to 45 (l45) 0.36 0.26–0.43 proportion V
Life expectancy at 45 20.7 12–24 further years Survivors of childhood often reach old age V
Adult hazard, ages 15–40 0.01–0.02 flat annual q Slope indistinguishable from zero V
Mortality-rate doubling time, 40+ 8–10 7–10 years Gompertz phase V
Cause of death: illness / violence / degenerative 70 / 20 / 9 % Whole cross-cultural sample V
Early Neolithic e0 (Vedrovice, LBK) 27.6 years Skeletal; ageing bias suspected V
TFR, foragers 5.6 <4–8 births/woman 12 populations V
TFR, horticulturalists 5.4 births/woman Not significantly different from foragers (p=0.8) V
TFR, intensive agriculturalists 6.6 ±0.3 SE births/woman The only significant subsistence break V
Interbirth interval, foragers 3.3 2.3–5.4 years V
Interbirth interval, sedentary horticulturalists 30.7 SD 10.6 months Tsimane V
Age at first birth 19.7 15.3–22.8 years Forager mean V
Age at last birth 39.0 26–42 years Forager mean V
Weaning age 24–48 12–72 months Foragers V
Sedentism → fertility (Agta, within-population) 7.7 vs 6.6 +16.7% TFR Settled vs mobile V
Sedentism → child mortality +63 0.93 vs 0.57 deaths/mother Same study — a quantity/quality trade V
Neolithic Demographic Transition +2 contested births/woman Driven by 3–4 month earlier weaning V, C
Steady-state agrarian growth 0.1–0.2 %/year Pre-industrial equilibrium V
Mean experienced band size 28.2 adults 32 societies, n=5,067 V
Adult primary kin per band 1.8 0.45–5.27 persons Only 7% of co-resident adults V
Household size, foragers ~5 2.9–7.7 persons V
Naroll floor-area constant 10 9–10 m²/person 18 societies; author called it "very rough" V
Late Natufian community 59 largest under 50 persons 0.2 ha V
PPNA community 332 18–735 persons 1.0 ha V
Çatalhöyük peak (revised 2024) 600–800 vs 3,500–8,000 older persons Revised down ~5× V

C-1. The founding population sits below every modelled threshold#

This is a structural finding. Read it together with C-4, which supplies the historical counterweight: real colonies of this size have occasionally survived, so the correct conclusion is marginal, not impossible — and what separates the survivors is network connection rather than headcount.

Every threshold estimated by every method sits above 18.

Method Threshold Finding
Demographic ABM (White 2017, ~5,000 runs/condition, 400-year horizon) 40–140 by marriage rules; 150 for near-certainty "Very small populations (i.e., less than 40 people) were not viable over a 400-year period"
Marriage-rule sensitivity ×2.25–3.25 Eight marriage divisions raise the requirement to 100–140; a simple incest taboo costs only ~10 persons
Closed-system Monte Carlo (Marin & Beluffi 2018) 98 for certainty 50 people went extinct in 50 ± 15% of runs when inbreeding was forbidden
Endogamous-population simulation (MacCluer & Dyke 1976) <300 monogamous; 50–100 with polygyny
Voyage microsimulation (Moore 2001) 150–180 "Populations varying in initial size from 4 to 60 persons invariably went extinct" at low fertility
Genetic (Frankham et al. 2014) Ne ≥ 100 "Ne ≤ 100 indicates … serious genetic threats after 5 or more generations"

The genetics are unambiguous at this size. Twelve breeding adults give Ne ≤ 12 — realistically 8–10 once family-size variance is included — so inbreeding accumulates at ΔF = 4.2% per generation, four times the tolerance the 50/500 rule was built around:

Generation F at Ne=12 Comparison
2 0.082 Already above first-cousin offspring (0.0625)
3 0.120 Approaching double-first-cousin
5 0.192 Beyond uncle–niece
10 0.347 65% of original heterozygosity remains

The founding event is mild — it retains 95.8% of heterozygosity. Sustained small size, not the bottleneck, is what does the damage.

The single best empirical anchor is the Raute of Nepal: a genuinely isolated population of 150 people — more than eight times December's settlement — with an explicit three-clan exogamy system deliberately optimized against inbreeding. Result: mean spousal relatedness r = 0.124 and F_ROH = 0.226, meaning 22.6% of the genome is autozygous, approaching full-sibling level. Their own analysis concludes that "owing to overall limited genetic variation, high expected offspring inbreeding is observed regardless of the simulated mating system." The gradient against better-connected groups is stark:

Population Mating-network size Mean spousal relatedness
Mbendjele BaYaka >30,000 regional 0.0053
Palanan Agta ~1,000 local / ~10,000 regional 0.0175
Raute 150, isolated 0.1240

C-2. What actually makes a small settlement viable: network membership#

The correction the second pass forced, and it improves the design rather than constraining it:

A forager band of 28 adults sustains a mostly-unrelated composition only because it continuously exchanges members with a wider network. The viable unit is the network, not the settlement.

The evidence is direct. Across 32 foraging societies, co-resident adult primary kin number only 1.8 per band — 7% of adults — and about a quarter of band members share no known genealogical or marriage tie with any given person at all. Among 80 Agta marriages in a ~270-adult network, 78 had no known shared ancestry. Studying two Pumé groups of ~11 adult males and ~11 adult females each — almost exactly December's scale — researchers concluded plainly that "in small populations, looking outside of one's local group is necessary to find a mate."

A methodological caveat that sharpens this. Figures drawn from ethnographic band size — the "magic 25," Dunbar's 150 — are not closed-population numbers. Those bands were never isolated; they maintained systematic mate exchange with adjacent bands, which is the entire reason a 28-adult band can be composed mostly of non-relatives. Citing band size as evidence that a small group is self-sufficient inverts the finding. The relevant quantity is always the network, never the camp.

Required network size: ~150–500 people, with ~175 the smallest well-supported value. This is where four independent estimates overlap: White's ABM (150), Wobst's mating-network simulation (79–332 raw, re-expressed as 175–475 under spatial assumptions), Birdsell's dialectal tribe (500 central tendency, but explicitly "commonly including groups of less than 200"), and the empirical forager hierarchy whose periodic-aggregation tier sits at ~165 and regional tier at ~839.

Required in-migration: roughly 1–3 exogamous marriages per generation, i.e. on the order of 15–30% of unions bringing in an outsider — comfortably within the ethnographic range.

But gene flow only rescues a population that is already big enough. The clearest demonstration comes from population viability analysis: at 12 breeding females with a realistic inbreeding load, one migrant per generation lowered 50-year extinction probability only from 100% to 97%, and three to four migrants gave no further improvement. At 48 breeding females, two migrants dropped it to 5%. There is a floor below which connection does not help.

C-3. Founder effects, and why "purging" does not rescue December#

A natural objection to §C-1 is that small populations purge their deleterious recessives — selection exposes them in homozygotes and removes them, so the inbreeding cost should be self-limiting. This is exactly the argument made against Frankham's revised thresholds. The evidence says purging is real but weak, acts only on the extreme lethal tail, and at n≈18 is dwarfed by drift.

Finding Evidence Class
Hutterite founder recessive lethals: 57% lost by 1950 — but at a rate "almost the same as for neutral variants" Loss was drift, not selection V
Direct human test of mating practices "Purging by non-random mating has low efficiency and different mating practices do not lead to different mutational loads" V
Actively consanguineous cohort Homozygous-knockout deficit only 13.7% (95% CI 8–20) — ~86% still occur V
Greenlandic Inuit, Ne <300 for 15,000+ years ≥20% MORE recessive load, not less V
Additive load Demography-proof; the two effects exactly cancel V
Counter-evidence Modest purging signal in recessive disease genes; ancestral Ne predicts persistence better than current Ne V, C

The Lord Howe Island stick insect, rebuilt from two mating pairs, is the cleanest illustration of the limit: stop codons were depleted, but "moderate- and low-impact mutations escape this process and may even fix." Purging removes the lethals and leaves everything else.

One result deserves to drive December's design directly. Extinction time for a population crashed to K=25–50 depends on its ancestral effective size, not its current one: 474 generations if the ancestral K was 1,000, but only 70 generations if the ancestral K was 15,000. A group that has always been small carries a purged, survivable load; a group that has recently become small carries the full reservoir of a large population and is far more fragile. December's founders left an established community — they are the second case.

Founding load, derived from verified rates (this arithmetic is the audit's, not a cited figure):

Quantity Value Basis
Recessive lethals per haploid genome 0.29 (95% CrI 0.10–0.84) Hutterite founder analysis
Distinct recessive lethal alleles entering an 18-founder pool ~10 36 haploid genomes × 0.29
Recessive LoF lethal equivalents per individual 1.6 Two independent studies converge
F by generation 10 at Ne ≈ 12–15, no immigration 0.30–0.35 ΔF = 1/(2Ne)

That projected F is roughly twice the mean autozygosity of UK Biobank's "extreme inbreeding" individuals (FROH 0.172, prevalence ~1 in 3,650), for whom the measured dose-response is:

Trait Effect at FROH ≈ 0.17
Peak expiratory flow −0.651 SD
Fluid intelligence −0.570 SD
Height −0.404 SD
Educational attainment −0.260 SD
Fertility RR ≈ 1.54 reduced

This is the best direct human dose-response available, and it points at the right mechanism for December: the cost lands on capacity and fecundity, not on dramatic death.

Founder contribution is severely unequal, which makes effective size far smaller than founder count. This is the most directly useful modelling finding in this subsection:

Population Nominal founders Effective concentration Class
Old Order Amish 554 in the 14-generation pedigree 128 account for >95% of the gene pool; 16 account for 50% V
Québec ~8,500 French settlers, 1608–1760 15% of founders explain 90% of genetic contribution V
Ashkenazi Effective founders ≈350 (250–420), 25–32 generations ago V
Hutterites 64 founders, 1,623-member pedigree Mean F = 0.034 V

December should not assume its twelve founders contribute equally. Differential reproductive success concentrates ancestry fast, so realized Ne will run below the 8–10 already assumed in §C-1.

The Hutterite fertility finding is directly implementable. Inbreeding significantly lengthened interbirth intervals (p=0.024) and time to conception (p=0.010), but produced no increase in fetal loss and no reduction in completed family size — reproductive compensation absorbed the cost. So December's inbreeding penalty should appear as slower reproduction under a longer shadow, not as dramatic infant death. That is both better sourced and more interesting than a mortality multiplier.

The best real-world anchor at December's exact scale is Tristan da Cunha: 15 settlers in 1816 (7 women, 8 men), 28 founders in total. The community survived to the present — but carries a 36% asthma prevalence today, alongside documented founder-effect retinitis pigmentosa. That is the honest picture of an 18-person founding at 200 years: persistence is possible, and it is not free.

For scale, human long-term effective population size sits at ~13,500 (ancestral non-African) with an out-of-Africa minimum around 1,200 at 40–20 kya. Even humanity's tightest documented bottleneck was roughly seventy times December's settlement.

C-4. The historical counterweight — small colonies that did survive#

Intellectual honesty requires reporting evidence that cuts against §C-1. The simulation and conservation-genetics literature says populations under 40 are non-viable. Actual history contains small founding groups that persisted anyway, and the audit's earlier flat statement that "18 people cannot persist" was too strong.

Case Founding party Outcome Class
Pitcairn, 1790 27 — 9 mutineers, 6 Polynesian men, 11 Polynesian women, 1 infant Persisted 230+ years. Peak 233 (1937); ~35 today. Violent early years left two adult males V
Tristan da Cunha, 1816 ~28 founders; 15 settlers (7 women, 8 men) Persisted. Carries 36% asthma prevalence and founder-effect retinitis pigmentosa V
Palmerston Island, 1863 4 effective founders — one man and three Polynesian wives, two of them cousins Persisted. 23 children, 134 grandchildren, >1,000 descendants by 1973 V
Pingelap, ~1775 ~20 typhoon survivors Persisted (~250 today) with ~10% achromatopsia and ~30% carriers V
Rapa Nui, 1877 Fell to ~110 (1892 census: 101 people, 12 adult men) Recovered. A documented real-world floor from which a population rebuilt V
Roanoke, 1587 117 Vanished V
Saint Croix Island, 1604 79 → 44 44% mortality in one winter (scurvy); abandoned V
Jamestown "Starving Time," 1609–10 214 → 60 72% mortality; colony voted to abandon, reprieved by a relief fleet V
Norse Greenland 300–500 landing; peak ~2,000 Extinct by ~1450 (latest ¹⁴C: AD 1430 ± 15) V
Henderson & Pitcairn (Polynesian) Abandoned after ~600 years of continuous occupation V

The corrected claim: an 18-person founding is marginal, not impossible. It sits below every modelled threshold, and the historical cases that survived did so while (a) retaining outside contact and periodic in-migration, and (b) carrying a permanent genetic cost — Tristan da Cunha's 36% asthma rate is what an unrescued founder effect looks like two centuries on.

Four historical mechanisms December should model#

These are the specific failure and survival modes the record actually documents. Each is more interesting than "the population was too small."

1. Founding sex ratio drives violence, and violence erases lineages. Pitcairn landed with 15 men and 12 women — a male surplus that the record ties directly to the killings. Within ten years, seven of nine mutineers and every one of the six Polynesian men were dead; in 1800 the colony held one adult man, nine women, and nineteen children. The genetic signature is unambiguous two centuries later: of 223 Norfolk Island male descendants, zero carry a Polynesian Y chromosome, while 40.4% of maternal lineages are Polynesian. One founding group's entire male line was annihilated by social conflict that began as a partner shortage.

That is a complete causal cascade — initial sex ratio → competition → homicide → permanent loss of half the founding ancestry — running on exactly the mechanisms December already models. It is the single best argument that this scale produces interesting history rather than merely fragile history.

Note also the recovery: Pitcairn went from 27 to 193 in 66 years at 3.0%/year growth, with infant mortality of only 5.5% and marital fertility above Hutterite levels. Small populations are not merely fragile; they are volatile in both directions.

2. Correlated risk can remove a whole cohort in a day. In 1885 Tristan da Cunha lost fifteen men — roughly 79% of its adult males — in a single boat, sent together to trade with a passing ship. The settlement never recovered its former size. December's activity scheduler should therefore treat co-location of a demographic cohort as a modelled hazard: a hunting party, a raid, a trading voyage, or a construction accident that puts most able adults in one place at one time is a single point of failure, and at n=18 it is an extinction-class event.

3. A cause can be cured and the collapse still arrive. St Kilda lost 45–69% of newborns per decade to neonatal tetanus for roughly 150 years, arising from a local umbilical-care practice rather than from isolation or inbreeding. The practice stopped and infant deaths reached zero by the 1920s — and the island was evacuated in 1930 anyway, because the age-structure hole left by five generations of lost children could not be refilled once the young adults emigrated. Demographic damage persists long after its cause is removed. A simulation that lets a settlement bounce back immediately once a hazard is fixed will miss this entirely; the lag is the story.

4. Exchange networks carried partners, not just goods. The prehistoric Polynesian colonies on Pitcairn and Henderson traded basalt, volcanic glass, pearl shell, and red feathers along a network centred on Mangareva — and the archaeological account is explicit that marriage partners were exchanged along the same routes. When Mangareva withdrew from the network in the 16th century, the satellites lost material supply and gene flow together, and both colonies died after five centuries of successful occupation. This is precisely §C-2's finding, observed rather than modelled.

5. Isolated communities invent institutions to manage their own genetics — and this is December's thesis in miniature. Palmerston Island was founded in 1863 by four effective people: one man and three Polynesian wives, two of whom were cousins. It produced over a thousand descendants. It did so by organising itself into three exogamous branches with intra-branch marriage prohibited — an invented kinship rule that structured mating away from the closest available unions.

Tristan da Cunha shows the same behaviour without the formal rule: a measured heterozygote excess, with a "pattern of homozygote deficiency suggestive of avoidance of close matings."

This is precisely the emergence December exists to produce — a material constraint (a shrinking pool of unrelated partners) generating a durable institution (an exogamy rule) that residents create, transmit, and enforce. It should therefore be available to agents, not imposed by the kernel. The kernel supplies the kinship graph and the consequences; whether residents notice the problem, propose a rule, get it adopted, and keep it enforced across generations is exactly the kind of question the project is built to answer. If a December settlement independently invents an exogamy institution under demographic pressure, that is a stronger result than any governance outcome currently listed in the success criteria.

A precision point on where the fitness cost lands. Couples who are cousins show no net fertility reduction — reproductive compensation absorbs it, and the Hutterite pattern (longer interbirth intervals, longer time to conception, no fetal-loss increase, no change in completed family size) is the mechanism. But individuals who are themselves inbred do show reduced fertility. The cost therefore lands one generation downstream of the consanguineous union, not on the couple that made it. That lag is both well-evidenced and dramatically more interesting to simulate than an immediate penalty, because it means the consequences of a marriage rule surface only after the people who set it are gone.

The recurring killer is network severance, not headcount. This is the single most useful pattern in the historical record and it converges exactly on §C-2:

  • Henderson Island and Polynesian Pitcairn sustained themselves for centuries and then died as dependent nodes whose supply network failed — both were satellites of Mangareva, and when Mangareva's network collapsed they could not stand alone.
  • Norse Greenland is structurally the same failure: regular Norwegian shipping ceased after the 1370s, and the ivory trade that had given Greenland a near-monopoly in western Europe collapsed when North Atlantic commerce shifted to bulk dried fish. Researchers call it a "rigidity trap."
  • Lynnerup's arithmetic on the Western Settlement is the sharpest version: at 600–800 people, 8–13 net emigrants per year — roughly one household — empties it in 200 years with no catastrophe at all. A settlement near the viability floor does not need a disaster; it needs only a slow leak.

A caution against overstating the genetics. The Norse Greenland "degeneration" hypothesis — small stature, inbred skulls — originated in 1920s racial anthropology, and modern reanalysis rejects it: the notorious "inbred monstrosity" skull was re-diagnosed as acromegaly, the claimed 6.5% stature decline "cannot be proven" because different regression equations were used across sites, and the definitive biological-anthropological study explicitly lists degeneration among the causes that can be rejected. What the skeletal record does show is modest: life expectancy at 20 fell ~1.5–3 years between early and late sites, concentrated in young adult females — and, tellingly, a possible decline in the sex ratio, since in a population this small "a chance series of death events may lead to extinction."

That last point is the real lesson. Small populations die from mate-availability failures and chance sex-ratio skews long before they die from inbreeding depression — which is precisely why §C-3's requirement for an explicit mate-availability model is the load-bearing one, not the inbreeding coefficient.

C-5. Consequences for the plan — revised#

  1. The neighboring aggregate group is mandatory at Phase 3, and the exchange and migration path must be validated before the raiding path is built. A boundary world introduced only as a threat would make the neighbors a hazard generator, which 07 forbids.
  2. One neighbor of similar size is not sufficient. Two settlements of 18 give a combined pool of ~36 — roughly a quarter of the minimum viable network. The scenario needs either a substantially larger neighbor, several neighbors, or an episodic regional aggregation connecting the valley to a wider pool. The last option is ethnographically standard (the periodic-aggregation tier of ~165) and is the recommended design: a seasonal gathering that residents travel to, which supplies partners, news, trade, disease, and diplomacy in one mechanic.
  3. Marriage rules are the dominant lever on viability — tightening them raises the required population 2.25–3.25×, far more than any other factor. December's kinship model is therefore not flavor; it is the parameter that most determines whether the world persists. It needs its own sensitivity sweep.
  4. A mate-availability model is mandatory. A demographic simulation without explicit partner matching will badly understate extinction risk — the audit's own quick Monte Carlo using verified vital rates but no marriage rules returned only ~1% extinction over 200 years, which is wrong for exactly this reason. Do not build vital rates without pairing.
  5. Fertility must not encode the wrong story. Foragers and horticulturalists have statistically indistinguishable TFR (5.6 vs 5.4, p=0.8); only intensive agriculture adds ~1 birth. But sedentism within a population raises fertility ~17% while raising child mortality ~63% — a quantity/quality trade, not a free gain. December's settlement is sedentary and pre-plough, so the horticulturalist band (TFR 6.0–6.5, IBI 30–34 months) is the right default.
  6. Null-model comparison remains mandatory (09).

D. Disease#

This section invalidated part of the original design and is unchanged from the first pass. Full sourcing in 13.

D-1. Critical community size#

Parameter Central Range Unit Context Class
CCS, measles (Bartlett) 275,000 250,000–300,000 persons US/UK cities, pre-vaccine V
CCS, measles (Black, islands) 400,000 300,000–500,000 persons Transmission broke in every community below 500,000 V
CCS, pertussis ~390,000 387,000–1,460,000 persons V
Crowd-disease population floor "several hundred thousand" persons Wolfe, Dunavan & Diamond 2007 V

December's population is four orders of magnitude below the measles CCS. No acute, directly transmitted, immunizing infection can be endemic in eighteen people. This is arithmetic over a seventy-year literature, mechanistically confirmed. An 18-person group produces roughly one birth every 1.5–3 years while measles needs new susceptibles on a two-week timescale.

D-2. The five-mechanism health model#

A single generic SEIR pathogen will either fade out immediately and contribute nothing, or — if tuned until epidemics appear at a satisfying frequency — encode a rate that cannot exist. That would be risk R-07 in epidemiological costume.

Mechanism Persists because Role in December
Introduced acute epidemics It does not — burns through and ends Rare punctuated shocks tied to contact events
Environmentally transmitted Environmental reservoir; CCS does not apply The persistent, self-generated disease pressure
Zoonoses with animal reservoirs Animal reservoir Rodent-borne risk from the settlement's own grain stores
Chronic and latent infection The host is the reservoir Background burden at any population size
Helminths and parasites Faecal-oral and environmental cycling Emergent consequence of sedentism and sanitation

D-3. Parameters and the finite-population outbreak requirement#

Parameter Central Range Unit Context Class
Contact-epidemic mortality 18 (median) <1–97 % of group 117 epidemics, 59 Amazonian societies V
Median affected population 180 persons Closest scale analogue to December V
Inter-epidemic period 7 rises with time since contact years V
Epidemic causes measles 37, influenza 25, malaria 13 % of epidemics V
Measles attack rate, virgin soil 77.5 % Faroe Islands 1846 V
Measles CFR, adequate care 2.8 % of cases Faroe Islands 1846 V
Measles mortality, care collapse 22 20–25 % of population Fiji 1875 — same pathogen, 10× the deaths V
TB lifetime reactivation 10 5–15 % of latent Elevated by malnutrition V
Helminth prevalence, foragers vs farmers 6 vs 76 % Ascaris Same ecosystem; sedentism is the variable V
Child deaths attributable to undernutrition 52.5 44.8–60.7 % V

R₀ is not a pathogen constant — measles estimates span 3.7–203.3 because R₀ is a function of social organization. Published values are upper bounds, not inputs; model contact structure and let the effective reproduction number emerge.

Introduced acute outbreaks at n=18 should often show strong fade-out-versus-large-outbreak behavior across ensembles. Under a simple early branching-process approximation with one index case, extinction probability is approximately 1/R₀; the final-size values below are large-population approximations, not exact invariants for eighteen heterogeneous people:

R₀ Final attack rate P(fade-out)
1.5 58% 67%
2.0 80% 50%
3.0 94% 33%
5.0 99% 20%
12.0 ~100% 8%

Many introductions should fizzle and some should affect much of the group; intermediate outbreaks remain possible in a finite, structured contact network. Validate the full stochastic distribution rather than only its mean, and do not mark an intermediate outcome as a bug solely because it differs from this approximation.


E. Violence — the range spans an order of magnitude and the dispute is substantially definitional#

This is the section where December must be most careful, because the temptation to pick a number is strongest and the literature least supports doing so.

Source Statistic Value Sample Class
Keeley 1996 % of all deaths from war 7–40 9 archaeological cases V, C
Bowles 2009 Fraction of adult mortality due to war (δ) mean 0.14, median 0.12 15 archaeological + 8 ethnographic V, C
Bowles 2009, range δ 0.00–0.46 Gobero 0.00 → Jebel Sahaba 0.46 V
Pinker 2011 % deaths, prehistoric sites mean 15, range 0–60 21 cases compiled from Keeley + Bowles V, C
Pinker 2011 Annual war deaths 524 per 100,000 Non-state societies V, C
Gat 2006/2015 Violent death rate ~25% of adult males, ~15% of adults HGs + pre-state horticulturalists V, C
Fry & Söderberg 2013 Lethal aggression events 148 events; median 4/society; range 0–69 21 mobile forager band societies V
Fry & Söderberg 2013 Events classified intergroup 33.8% overall; 15.2% excluding Tiwi V
Ferguson 2013 Pinker's list after audit 21 cases → 14 Duplicates, single deaths, misreadings removed V
Jurmain, via Ferguson Jebel Sahaba recount 9.8% vs Bowles's 46% 4 of 41 complete skeletons V
Meijer 2024 Prehistoric HG lethal violence notes ~2–3% of deaths, declines to endorse Global archaeological review V

E-1. Why the numbers cannot simply be averaged#

The two flagship papers do not report the same statistic. Bowles reports war deaths as a fraction of all deaths. Fry & Söderberg report counts of lethal events classified by motive and never compute a mortality rate at all. They are routinely presented as opposing estimates of one quantity. They are not.

Six of Bowles's eight ethnographic societies also appear in Fry & Söderberg's twenty-one — the same societies, opposite conclusions. That difference is definitional, not empirical. Bowles defines war as any coalitional lethal action across group boundaries, explicitly including revenge killings; Fry & Söderberg follow Kelly in requiring social substitutability (any member of the offending group being a legitimate target), which reclassifies much of the same behavior as feud or homicide.

The denominators differ too: Keeley counts all individuals, Bowles counts adults only, Gat's 25% is of adult males, and frequently cited trauma figures (57.3% of Australian crania; 10.91% Neolithic cranial trauma) count injuries, mostly healed, not deaths. Four different quantities, routinely juxtaposed.

Both headline results are leveraged on outliers. A single society, the Tiwi, supplies 69 of Fry & Söderberg's 148 events and 38 of their 50 intergroup events; removing it halves the mean and cuts the range from 0–69 to 0–15. On the other side, Ferguson's audit removed 7 of Pinker's 21 cases as duplicated sites, single deaths, or misread evidence.

Sample composition is a deliberate choice on both sides. Bowles knowingly included sedentary hunter-gatherers and seasonal forager-horticulturalists (3 of his 8); Fry & Söderberg deliberately excluded them. Fry's earlier work found 62% of mobile forager band societies non-warring while all complex and equestrian forager societies had war. Most of the famous archaeological massacre sites — Crow Creek, Talheim, Schöneck-Kilianstädten, Asparn/Schletz, Potočani — are farming societies, not foragers, and bear on a different question than the one they are cited for.

Jebel Sahaba is the cautionary tale. The single most-cited site has been reported at 46%, 40.7%, and 9.8% of deaths, and a 2021 microscopic reanalysis of all 61 individuals found more total violence than previously documented while concluding the evidence "dismisses the hypothesis that Jebel Sahaba reflects a single warfare event," supporting recurrent raiding instead. The same site supports opposite headlines depending on whether you count trauma or count events.

E-2. What December must do#

Treat conflict frequency as an output, not a direct input knob. Sweep mechanisms that could produce it—resource pressure, mobility, group boundaries, threat sensitivity, dominance motives, institutional effectiveness, retaliation, logistics, and provider refusal—and compare the resulting outputs with broad contested reference ranges. Concretely:

  1. Pick and publish a definition before running anything. December must state whether it counts coalitional cross-boundary killing (Bowles) or requires social substitutability (Kelly/Fry), and it must count events and mortality fractions separately, with explicit denominators. Most of the published disagreement is exactly this choice, so making it silently would be inheriting a position without argument.
  2. Sweep mechanisms, and report outcome distributions. The literature spans roughly 2% to 25% of deaths depending on definition, denominator, and society type. December must not choose a target rate and tune until it appears.
  3. Claim nothing about human nature. The prohibition in 09 is reinforced: because the empirical range is an order of magnitude wide and partly definitional, December's simulated conflict rate carries no evidentiary weight in either direction. It cannot show that violence is natural, and it cannot show that peace is.
  4. What December can honestly demonstrate is the structural finding both sides accept: that sedentism, population density, storable and defensible resources, and social segmentation are associated with more organized violence. That is common ground across Kelly, Fry, Ferguson, Meijer, and Gat, and it is a mechanism December actually models. Showing conflict emerging from those pressures is a defensible result; showing it at a particular rate is not.
  5. Risks R-11 (conflict as entertainment) and R-07 (catastrophe machine) both bite here, and the refusal-rate requirement in 16 §6 applies to every conflict statistic.

The most honest summary available comes from Ferguson, a partisan of the low side, in two consecutive sentences: anyone believing violence began with colonialism, the state, or agriculture is proven wrong — and equally, anyone believing all human societies were plagued by war is proven wrong. Recent reviews (Meijer 2024, Glowacki 2023) conclude the debate is at an impasse that the archaeological record probably cannot resolve, and that the true finding is the variance itself.


F. Movement#

Parameter Central Range Unit Context Class
Tobler's hiking function `W = 6·exp(−3.5· S + 0.05 )` km/h
Tobler, flat ground 5.04 km/h D
Tobler, maximum 6.00 at S = −0.05 km/h Slight downhill is fastest D
Tobler, off-path multiplier ×0.6 Tobler's own text V
Tobler vs measured −34 % (predicts too fast) 200 GPS-tracked walkers V
Tobler percentile ~5th percentile Against 29,928 Strava users — it is a slow walker V
Naismith's rule 1 h/5 km + 1 h/600 m ascent Sample size: n=1, a club trip report V, C
Langmuir gentle descent (5–12°) −10 min/300 m Subtract V
Langmuir steep descent (>12°) +10 min/300 m Add; creates a discontinuity at 12° V
Sustainable load, non-specialist 25 20–30 % body mass Multi-day walking V
Fighting / approach-march load 30 / 45 % body mass Military doctrine V
"Free ride" below 20% body mass 82% of 45 studies fail to replicate it V, C
Professional porter load 85 80–200 % body mass Lifelong specialists; −20% metabolic cost V
Loaded travel, sustainable 26 20–32 km/day Roads, daylight V
Forced march, 24 / 48 / 72 h 56 / 96 / 128 km cumulative Note the sublinearity V
Hadza daily distance M 11.4–12.9 / F 5.8–7.6 km/day GPS, two studies V
Hadza foraging trip M 8.3 / F 5.5 km S
Max comfortable daily round trip 25 20–30 km Hunters, many habitats S
Foraging radius before camp move 6 1.5–10 km Function of camp-move cost S
Load speed penalty −28 % (1.25 → 0.9 m/s) !Kung women carrying ~30% body mass V
Walking speed, dense forest 1.67 km/h V

Design note F-1 — the depletion halo is a documented, directly implementable mechanic#

The !Kung record gives December its ecology-and-labor coupling almost for free. A camp exhausts mongongo nuts within 1.5 km in week 1, 3 km in week 2, and 5 km in week 3; measured daily round trips rose from 9–14 km in June to 19 km by August. Foraging radius before a camp move is not a constant but a function of move cost — a two-hour camp breakdown makes it worth exhausting resources within ~6 km, while a half-hour breakdown makes moving worthwhile at 1.5 km.

Combined with Tobler's terrain-sensitive travel times, this produces the continuously rising subsistence cost that 03 wants from its ecology — not a stepwise "resource depleted" flag, but a slowly tightening squeeze that residents can perceive, argue about, and respond to by moving, intensifying, or fighting. It should be an explicit Gate 1 pattern target.

Two implementation cautions: Tobler's S is rise/run — published archaeology has got this wrong by a factor of 100 — and Tobler predicts roughly a 5th-percentile walker, ~34% faster than GPS-measured off-road speeds, so it is a slow baseline rather than a median one. The forced-march sublinearity (56 km on day 1, only 40 more by day 2, only 32 more by day 3) is the most useful single fact for modelling multi-day journeys.


Gaps this registry does not close#

  1. No crop phenology parameters — stage durations, temperature and moisture thresholds, damage functions. Needed at Phase 1.
  2. No hydrology parameters — infiltration, runoff, baseflow recession, irrigation efficiency.
  3. No construction labor figures — person-days per structure type, the direct input to the project compiler in 05. The swidden labor breakdown in §B is the closest available analogue.
  4. No stone-tool forest clearance rate — sought and not found; use the swidden slashing-and-felling column (9–75 person-days/ha) as the stand-in.
  5. No tool wear or maintenance rates.
  6. No skill acquisition curves — the "diminishing returns" in 04 remains unparameterized.
  7. Fire spread parameters deferred to Phase 5, entirely open.
  8. Pre-modern road-network travel rates (Roman/medieval) — not obtainable in this pass.

Items marked N or S above are the ones to verify first. Every gap that ships uncalibrated becomes a registry row with source: NONE — DESIGNED, disclosed in the known-limit report required at canonical launch.

Research and hardeningwiki/16-cost-model-and-model-selection.md
3,481 words
Inside this chapter

16 — Bottom-Up Cost Model and Model Selection#

Status: added by the 2026-08-01 audit pass. All prices observed 2026-08-01 and will be stale within weeks — re-derive before committing a budget.

Why this document exists#

Risk R-04 ("token spend explodes") is rated High/High in 11, and the whole project is gated on "continuous operation within bounded cost." Yet the only number in any prior draft was an illustrative arithmetic exercise on someone else's published token totals. A project whose top operational risk has no bottom-up estimate cannot pass its own Gate 0.

This document builds the estimate from activation counts upward, prices it against the actual August 2026 market, and identifies which decisions dominate the bill.

The Agentopia reference point, restated correctly#

Prior drafts cited Agentopia's token volume without stating what it was. Verified against the paper (arXiv:2606.07513, Table 4):

  • 13,347 M input tokens, 352 M output tokens, 567 K LLM calls — these are averages across three simulation worlds, not a single run. Per-run input ranged 9,699–19,041 M.
  • The run covered 100 agents over 10 simulated years and took ~186 wall-clock hours.
  • The model was Qwen3.5-397B with Gemini 3 Flash as a fallback for invalid outputs — a detail that matters enormously and is discussed in §5 below.

Two derived figures are the useful ones:

Derived quantity Value Why it matters
Calls per agent per simulated day 1.55 Sparse activation is achievable at scale
Input tokens per call 23,540 Contexts are large; input dominates
Input : output ratio 38 : 1 Prompt caching is the primary cost lever, not output brevity
Simulated days per wall-clock hour 19.6 Throughput is not the binding constraint

The old draft's "$4.4k" substitution should be discarded. It priced someone else's workload at a rate for a different model and told us nothing about December.

1. Activation model#

December's cognition is denser per agent than Agentopia's (an explicit appraise/deliberate/review cycle rather than a weekly loop) but runs on far fewer agents. Central case, per simulated day:

Call class Callers Calls each Total calls Input each Output each
Tier B — appraisal, routine choice, short speech 12 full residents 12 144 4,500 300
Tier B — lightweight residents 6 dependants/elders 1 6 3,000 200
Tier C — consequential deliberation 12 full residents 1.5 18 12,000 900
Director / observer summarization 10 8,000 1,200
Total per simulated day 178 ≈ 971 K ≈ 73 K

Input : output ≈ 13 : 1.

Cross-check against Agentopia. December: 971 K input ÷ 18 residents ≈ 54 K input tokens per resident per simulated day. Agentopia: 1.55 calls × 23.5 K ≈ 36 K. This is a sanity check, not independent validation: both are speculative LLM-society workloads with different cognition loops. Until a four-resident prototype measures activation, context, repair, refusal, and cache-hit rates, use the wide envelope below plus a separate experiment reserve.

Uncertainty is large and asymmetric. Activation rate and context size each plausibly vary by 2× in either direction, so the honest envelope is 0.25× to 4× the central case. All figures below should be read with that band.

2. Price landscape, August 2026#

Model Input $/M Output $/M Notes
Zhipu GLM-4.7-Flash 0.06 0.40
OpenAI gpt-5-nano 0.05 0.40
Qwen3.5-Flash 0.10 0.40
DeepSeek v4-flash 0.14 0.28 Cache hits at $0.0028 — a 50× read discount
OpenAI GPT-5.6 Luna 1.00 6.00 Official GPT-5.6 family pricing; 30-minute minimum cache life
Google Gemini 2.5 Flash-Lite 0.10 0.40
Google Gemini 3.5 Flash-Lite 0.30 2.50
MiniMax M3 0.30 1.20 ≤512 K tier; official page labels this Permanent 50% off, not a temporary promotion
MiniMax M2.7 0.30 1.20 Not marked promotional
Anthropic Claude Haiku 4.5 1.00 5.00
Anthropic Claude Sonnet 5 3.00 15.00 Introductory 2.00 / 10.00 through 2026-08-31

Corrections: M2.5 is legacy; M3 is the current flagship (1 M context) and M2.7 the mid-line workhorse. MiniMax's official page describes M3's displayed reduction as permanent, so the first audit was wrong to call it temporary and to advise budgeting at a higher "list" price. Prompts above 512 K cost twice the displayed standard tier.

Re-verified 2026-08-01 (third check): GPT-5.6 Luna's standard tier is $0.20 input / $1.20 output per million (short context), with long context at $0.40/$1.80, batch and flex at $0.10/$0.60, and fast mode at $0.40/$2.40. The second audit's assertion of "$1/$6" does not correspond to any tier on the official pricing page and has been withdrawn. The row is restored below:

Model Input $/M Output $/M Notes
OpenAI gpt-5.6-luna 0.20 1.20 Standard, short context; batch/flex 0.10/0.60

This is the second time a pricing figure has moved on re-check. Treat every price in this document as a snapshot requiring verification before it is committed to a budget, and prefer fetching the provider page over trusting any summary — including this one.

3. Cost per simulated day and per month#

Two paces from 01: attended (1 real hour = 1 simulated day → 730 simulated days/month) and unattended (1 real hour = 3 simulated days → 2,190 simulated days/month).

The following planning scenario assumes 65% of input is a stable cacheable prefix read at roughly 0.1×. That is not provider-neutral: cache semantics, routing affinity, TTL, writes, and OpenRouter provider selection vary. Gate 2 replaces it with measured hit/write rates.

Model (single-tier) $/sim-day uncached $/sim-day cached $/month @ 730 $/month @ 2,190
gpt-5-nano 0.08 0.05 37 110
GLM-4.7-Flash 0.09 0.05 37 110
DeepSeek v4-flash 0.16 0.07 51 153
MiniMax M3 (current ≤512 K tier) 0.38 0.21 153 460
Gemini 3.5 Flash-Lite 0.47 0.27 197 591
Claude Haiku 4.5 1.34 0.77 562 1,686
Claude Sonnet 5 4.01 2.30 1,679 5,037

The recommended mixed-tier configuration — Tier B on a cheap model, Tier C on a strong one:

Component Model $/sim-day
Tier B (150 calls, 675 K in / 45 K out) DeepSeek v4-flash 0.05
Tier C (18 calls, 216 K in / 16 K out) Claude Sonnet 5 0.51
Director (10 calls) cheap tier 0.01
Total ≈ 0.57
@ 730 sim-days/month ≈ $416
@ 2,190 sim-days/month ≈ $1,248

4. What actually drives the bill#

Finding 1 — Tier C model choice dominates, not agent count. In the mixed configuration, Tier C is 10% of calls and 89% of cost. Doubling the resident population is cheaper than upgrading the Tier C model one tier. The decision-significance threshold in 04 is therefore the primary cost-control instrument in the entire system, and it deserves the engineering attention that would otherwise go to trimming prompts. It should be a live, tunable, monitored parameter with its own dashboard, not a constant buried in config.

Finding 2 — sparse activation and prompt caching are in direct conflict, and nobody noticed. This is the sharpest operational finding in this document.

Input is 93% of tokens, so caching the stable prefix (identity block, world rules, action schema, consolidated memory) is the difference between a $50/month world and a $500/month one. But cache entries expire on wall-clock TTLs:

Provider Cache TTL Cache read Cache write
DeepSeek hours to days ~0.02× free
OpenAI (GPT-5.6+) 30 min 0.1× 1.25×
Google Gemini 1 h explicit (storage-billed) 0.1× no premium
Anthropic 5 min default; 1 h option 0.1× 1.25× / 2×
MiniMax 5 min ~0.2× 1.25×

Meanwhile 07 makes activation deliberately sparse and irregular. At the attended pace, a resident activated ~13 times per simulated day is activated ~13 times per real hour — roughly one call every 4.6 real minutes, which straddles a 5-minute TTL. Every cache miss costs 10× the token price, so an architecture optimized for fewer calls can easily cost more than one optimized for cache locality.

Required responses:

  1. Cache-aware activation scheduling. When several residents are due for cognition within a short window, batch them so each resident's prefix is reused inside its TTL rather than scattering calls across the hour.
  2. Treat cache TTL as a model-selection criterion of the first rank. DeepSeek's multi-hour cache and OpenAI's 30-minute window are worth more to December than a modest per-token price advantage. This belongs in the model profile record in 08.
  3. Measure cache hit rate per resident per tier as a primary operational metric, alongside cost. A falling hit rate is the leading indicator of a cost incident.
  4. Structure prompts prefix-stable. Identity and rules first, volatile observations last. A single early-token change invalidates the whole prefix.

Finding 3 — batch APIs are largely unavailable to a live world. Anthropic, OpenAI, Google, and Alibaba all offer ~50% batch discounts, but with turnarounds up to 24 hours. That is unusable for canonical cognition. It is usable, and should be used, for: shadow-world experiments, kernel ensembles, offline Tier D analysis, memory consolidation jobs, and model-admission suites. Those are exactly the workloads 09 says will run in the thousands, so the saving is real — just not on the canonical path.

Finding 4 — the accelerated experiments, not the canonical world, are the budget risk. 09 calls for "dozens/hundreds of multi-season runs with bounded LLM use." A single 100-simulated-day run at the central rate costs ~$57 in the mixed configuration; 200 such runs is ~$11,400 — an order of magnitude above a year of canonical operation. The experiment budget needs its own hard cap and its own approval, and should default to the cheapest tier plus batch pricing. The existing budget hierarchy lists a "shadow-world/experiment cap" but does not flag that it is the larger number.

5. Structured output: constrained decoding is uneven, so local validity is mandatory#

04 requires every model output to validate against a versioned schema, and 08's admission suite makes schema compliance the first gate. Strict schema-constrained decoding is uneven across providers. That does not make lower-cost models unusable: tool calls plus local validation, repair, and escalation can satisfy the transport contract. Strict JSON is also not semantic validity—a perfectly shaped command may still be irrational or infeasible.

Provider Strict JSON-Schema guarantee Notes
OpenAI Yes additionalProperties: false required; all fields required; recursion supported
Anthropic Yes No recursive schemas; no numeric/length constraints
Google Gemini Yes (structural) Supports recursion and numeric bounds
DeepSeek No — best-effort JSON mode only Docs warn output may be empty
Qwen No — best-effort JSON mode only
MiniMax No — undocumented, open feature requests Tool calling works; response_format unsupported

This constrains rather than eliminates MiniMax for routine cognition. “OpenAI-compatible” does not imply response_format compatibility. MiniMax’s current OpenAI-compatible documentation lists tools but not structured-response formatting, so conformance must be measured through tool-call framing and local validation before admission.

Note that Agentopia hit exactly this wall and solved it the same way we should: Qwen for volume, Gemini Flash as a fallback for invalid outputs. That is independent confirmation of the design.

Required architecture — the validity ladder. Every Tier B/C call resolves through:

  1. Constrained decoding where the provider supports it.
  2. Otherwise tool-call framing with a strict schema, which is better supported than response_format on most OpenAI-compatible endpoints.
  3. Local validation against the versioned schema — always, regardless of provider claims.
  4. One repair attempt on the same model with the validation error appended.
  5. Escalation to a schema-guaranteed model (Gemini Flash-Lite or gpt-5-nano) for the repair, not a retry on the same model. This is the Agentopia pattern and it is cheap: it only fires on failures.
  6. Deterministic Tier A fallback, recorded as a DecisionResolved event with resolution: fallback per 14.

Track schema-failure rate per model per decision type as an admission criterion and an ongoing metric. A model whose failure rate rises after a silent provider update is quarantined.

6. Model refusal is an unhandled failure mode with emergence consequences#

Not addressed anywhere in the prior plan. December asks models to act as residents who may raid, steal, deceive, threaten, withhold food from rivals, and make reproductive decisions. Commercial models sometimes decline such requests, break character to comment as an assistant, or soften an action.

This is worse than a nuisance — it is a systematic bias that silently invalidates emergence claims. If a provider's safety layer makes agents reluctant to escalate conflict, December will report that "peaceful institutions emerged from material conditions" when the true cause was reinforcement learning at the provider. The 09 emergence audit asks whether any prompt contained the outcome label beforehand; it must also ask whether the model declined the alternative.

Required:

  1. A refusal-rate benchmark in the model admission suite (08), covering the full action grammar including conflict, deception, theft, and household formation, framed as the analytical simulation content it is.
  2. Refusals are a first-class failure class, logged distinctly from schema failures and timeouts, and surfaced as an operational metric.
  3. Per-model refusal profiles are part of the model profile record, because a model swap that changes refusal behavior is a behavioral intervention on the world and must be recorded as such under the change-management rules in 08.
  4. Refusal rates are reported alongside any conflict-frequency finding. A conflict statistic without its refusal denominator is uninterpretable.
  5. Prefer models with steerable, documented behavior over the December action grammar even at a price premium, and prefer open-weight models where refusal behavior can be measured stably over time.

7. Storage and retention costs#

14 estimates 40–150 GB per real year of event log at the unattended pace under the batched-event policy, plus 1–5 MB per simulated day of prompt/response cache — which is canonical and must be backed up. At 2,190 simulated days/month that cache alone is roughly 26–130 GB/year.

Total canonical artifact growth: order 100–300 GB per real year, replicated. This is affordable on local disk and cheap in object storage, but it is not zero, and the monthly restore drill in 08 must be timed against these volumes rather than against an empty database.

7a. The actual budget: $2,000 MiniMax + $1,000 OpenRouter#

The owner holds $2,000 in MiniMax credits and $1,000 in OpenRouter credits. This is now the real constraint, and it answers open question 1 in 11 for the pre-revenue phase.

The structure of these credits matters more than the total, because the two pools are not interchangeable.

MiniMax $2,000 OpenRouter $1,000
Spendable on MiniMax models only ~any routed model
Strict JSON-schema decoding No — undocumented, open feature requests Yes, via Gemini / OpenAI models
Prompt-cache TTL 5 minutes — the shortest of any provider Varies by routed model
Cache read $0.06/M vs $0.30/M input — a 5× saving Varies
Batch discount None None
Model diversity One family Many families

Three consequences follow directly.

The pools are complementary by necessity, not merely by price. MiniMax cannot guarantee schema-valid output, so if it carries Tier B volume, the validity ladder in §5 is load-bearing from the first day of Phase 2 and its escalation target must live on OpenRouter. Conversely, OpenRouter is the only pool that can supply schema guarantees, Tier C reasoning, and the multi-family model assignment that mitigates behavioural homogeneity (R-31). Spending OpenRouter credits on bulk Tier B traffic would waste their only unique properties.

Neither pool offers a batch discount. §4 Finding 3 recommended batch pricing for the ensemble programme; that 50% saving is unavailable on these credits. Ensembles either pay full rate here or run against a direct provider account that has a batch API.

MiniMax's 5-minute TTL sits exactly where the token volume is. Cache-aware activation batching (§4 Finding 2) therefore becomes more important under this budget, not less — a hit is 5× cheaper than a miss on the dominant cost line.

What the credits buy#

Recommended allocation — MiniMax carries volume, OpenRouter carries correctness and judgment:

Component Pool / model $/sim-day
Tier B, 150 calls (675 K in / 45 K out) MiniMax M3 0.15 cached – 0.26 uncached
Director, 10 calls MiniMax M3 0.04
MiniMax subtotal 0.19 – 0.30
Tier C, 18 calls (216 K in / 16 K out) OpenRouter → Gemini 3.5 Flash-Lite 0.11
Schema repair escalation OpenRouter → gpt-5-nano fires only on failure
Combined ≈ 0.30 – 0.41

Horizon: roughly 8,000 simulated days — about 22 simulated years — with both pools exhausting at approximately the same time. The MiniMax pool alone yields ~6,700–10,500 sim-days; matching that with $1,000 of OpenRouter requires Tier C to stay near $0.125/sim-day, which Gemini 3.5 Flash-Lite ($0.105) hits almost exactly.

That balance is fragile in one direction. Tier C alternatives at $1,000:

Tier C model $/sim-day Sim-days from $1,000 Outcome
gpt-5-nano 0.02 ~58,000 OpenRouter never binds; weakest deliberation
Gemini 3.5 Flash-Lite 0.11 ~9,500 Balanced — both pools end together
Claude Haiku 4.5 0.30 ~3,400 OpenRouter dies first, ~$1,200 MiniMax stranded
Claude Sonnet 5 0.89 ~1,100 OpenRouter dies at 3 sim-years, ~$1,750 stranded

Choosing a Sonnet-class Tier C spends the OpenRouter pool seven times faster than the MiniMax pool and strands most of the larger credit. This is §4 Finding 1 restated in cash: the Tier C model is the budget.

What 22 simulated years is, and is not#

At the two documented paces: ~11 real months attended (24 sim-days/real-day), or ~3.7 real months unattended (72 sim-days/real-day).

Demographically it is exactly one generational transition. Founders age 22 years; children born at t=0 reach the ethnographic age at first birth (~19.7) just before the credits run out. That is enough to watch the kinship constraint in 15 §C arrive and force a response — the exogamy-institution question in 05 becomes live and testable. It is not enough for deep multigenerational history, inherited institutions across several generations, or the long ecological trends in 03.

The allocation conflict the owner must resolve#

§4 Finding 4 warned that the experiment programme, not the canonical world, is the larger budget line. Under these credits that is now concrete:

  • A canonical world to one generational transition: ~8,000 sim-days — the entire pool.
  • The ensemble programme in 09 ("dozens/hundreds of multi-season runs"): 200 runs × 100 sim-days = 20,000 sim-days — 2.5× the entire pool, with no batch discount available.

These credits cannot fund both. The choice is a research-programme decision, not an accounting one, and it belongs with the claims ladder in 18: a registered cohort study with controls is the scientific artifact, while a long canonical exhibit is the demonstration. Pass 2's separation of exhibit from cohort (P2-F12) is what makes the trade-off visible.

Recommended split, pending the owner's decision:

  1. Reserve ~$500 of OpenRouter for model admission, refusal benchmarking, homogeneity testing, and schema-failure measurement across families. These are prerequisites to trusting anything else and they are cheap.
  2. Spend MiniMax on Phase 2–3 development traffic, where volume is high, stakes are low, and schema failures are informative rather than costly.
  3. Do not start a long canonical run until Gate 3, when measured activation counts replace §1's estimates. Starting early burns the pool on a world whose cost model is still an estimate.

Phase 1 requires no LLM spend at all — the headless kernel has no cognition — so the credits remain untouched through the longest build phase and are not at risk from the schedule.

Uncertainty#

§1's activation model carries ±2× on both activation rate and context size, so the honest envelope on the 8,000-day figure is roughly 2,000–32,000 simulated days. Every number here is planning arithmetic over an unmeasured workload; Gate 2 replaces it with measurement. Track spend per simulated day from the first cognition call.

  1. Set the monthly cap before Phase 1, as 11 already requires. This document supplies a scenario—not a forecast: roughly $50–$1,300/month under current assumptions, with wider contingency until the four-resident prototype measures real behavior.
  2. Budget at the official current price plus contingency, and model known context-tier step changes.
  3. Separate and cap the experiment budget, which is likely the larger number.
  4. Instrument cost per simulated day as a first-class metric with an alert on trend, not just on absolute spend. Input-token growth per simulated day is the early warning that context assembly is regressing.
  5. Re-derive this entire document at Gate 1 with measured activation counts and context sizes replacing the estimates in §1. Everything here is an estimate whose purpose is to make the risk tractable, not to be right.
Research and hardeningwiki/17-initial-conditions-and-authorship.md
1,473 words
Inside this chapter

17 — Initial Conditions: Declared Authorship and Experimental Variation#

Status: added by the 2026-08-01 audit pass.

The gap this closes#

09's emergence audit asks: "Did any prompt, fixture, or director text contain the outcome label beforehand?" That is the right question, and the plan had no answer for the largest fixture of all — the founding world itself.

Nowhere did any prior document specify who writes the twelve residents' biographies, values, traits, skills, and relationships; who sets the initial stores, claims, and land quality; or where the founding charter comes from. Yet 03 already ships a scenario in which the founders "disagree about property, leadership, risk, and relations with a neighboring mobile band," claims over the best land are "ambiguous," and the charter conveniently leaves assembly procedure "unsettled."

Those are not neutral initial conditions. They are the plot, pre-installed. If the settlement later fractures over a property dispute and elects a leader to resolve it, December will have observed the thing it planted. The causal graph will be perfectly honest and completely uninformative, because the first cause lies outside the recorded history.

This is a Blocker-class confound for the project's central claim. It is also entirely fixable.

The authorship boundary#

Principle. Every initial condition is authored, including one produced by a generator: developers choose variables, distributions, correlations, exclusions, validation bands, and which seed becomes canonical. The honest objective is not an “un-authored” state; it is declared, reproducible, varied, and non-cherry-picked authorship.

Initial conditions are part of the experimental design. Research cohorts sample registered initial-condition families. The public canonical world may be selected for observer legibility, but if it is selected rather than randomly drawn it is labeled an exhibit and excluded from unbiased cohort claims.

Three consequences:

  1. Anything hand-authored in the founding state is a developer intervention at t=0 and must be visible as such in the audit lineage.
  2. Any initial condition that names or structurally guarantees an outcome later claimed as emergent is a confound. An outcome-word scan is useful but insufficient; review must also inspect proxy features and correlations that plant the same result without naming it.
  3. The founding state must be resampleable, so that ensembles vary initial conditions rather than treating one hand-built world as ground truth.

What must be removed from the current scenario#

The Founding Valley description in 03 must be rewritten to state material and structural facts only. Specifically:

Currently stated Problem Replacement
Founders "disagree about property, leadership, risk, and relations with the neighboring band" Pre-installs the four axes of every subsequent conflict State the material facts that could produce disagreement: unequal stores, overlapping use claims, differing household sizes and dependency ratios, different distances to water
Claims over the best land are "ambiguous" Pre-installs the dispute Record actual claims with actual evidence and actual overlaps; let ambiguity be a derived property of the claim ledger
Charter leaves assembly procedure "unsettled" Pre-installs the constitutional crisis Either the charter specifies a procedure or there is no assembly; "deliberately underspecified so it can be fought over" is authorship
"Neither utopian nor already doomed" Reasonable as a design intent, but unfalsifiable as a spec Express as a measurable initial condition: stores cover N days at current consumption, and the ensemble's 1-year survival rate under scripted policies falls in a declared band

The general rule: a founding fact is legitimate if it is a quantity, a location, a relationship with evidence, or a capability. It is illegitimate if it is a disposition toward a future conflict.

Traits and values are the hard case. Residents do need heterogeneous values — 07 correctly lists "heterogeneous but plausible goals/values" as a driver of possibility space. The distinction:

  • Permitted: sampling each resident's values from a declared distribution over the value vocabulary in 04 (reciprocity, autonomy, tradition, security, status, generosity, truthfulness), with the sampling seed and distribution recorded.
  • Prohibited: authoring a specific resident as "resents the steward" or "believes the upper field is hers by right," which is a grievance and a claim, not a value.

Biography is likewise permitted as history (where someone came from, what they know how to do, who they are related to, what they have done) and prohibited as foreshadowing.

Generation procedure#

The founding world is produced by a seeded generator, versioned like any other kernel component, and run before t=0.

world_seed
  → terrain and hydrology synthesis (elevation, soil, water, vegetation, resource stocks)
  → resource and structure placement
  → household composition (sizes, ages, kin links) sampled from the demographic ranges in [15]
  → per-resident skills, values, and biography sampled from declared distributions
  → initial stores, tools, claims, and the founding charter
  → validation pass
  → world manifest entry recording generator version, seed, and every distribution used

Rules:

  1. Seeded and reproducible. Same seed and generator version produce the same founding world, verified by the state-hash machinery in 14.
  2. No LLM in the canonical generator. Names, biography text, and flavor may be LLM-generated as a separate, cached, content-addressed step, but no LLM decides a quantity, relationship, claim, or capability. Prose describes the founding state; it never determines it.
  3. The generator is ablatable. Ensembles must be able to resample residents while holding terrain fixed, and vice versa, so 09 can attribute outcome variance to initial conditions versus dynamics. This is the only way to answer "would this have happened in a different founding world?"
  4. Validated before use. The founding state passes the same invariants as any other state: no negative stocks, consistent claims, reachable water, feasible travel times, conserved totals.
  5. Recorded as genesis events. The founding world enters the log as a genesis event set with actor: generator@version and authorization: scenario_config, so the audit trail begins before the first resident acts rather than at an unexplained initial snapshot.

The founding charter#

03's three consensual rules are a reasonable minimum institutional seed, but they must be justified rather than assumed, because 12 lists "elections, police, war, or religion installed as assumptions despite emergence claims" as an automatic High finding.

The defensible position: the founders left an existing community, so they carry inherited norms. That is a fact about their history, not a designed outcome. Therefore:

  • The charter is generated as inherited culture with recorded provenance, not hand-written to be interesting.
  • A no-charter condition belongs in the factorial design when the question concerns institutional emergence. It need not appear in every unrelated ensemble. If institutions form only with inherited rules, that is evidence that cultural inheritance is causal, not a void result; the reported claim must say so.
  • The charter's contents are swept, not fixed: charter-with-3-rules, empty charter, and a differently-seeded charter should all appear across runs.

Operationalizing the success criteria#

00 claims outcomes such as "at least three distinct governance forms arise across seeded runs." As written this is unfalsifiable — nothing defines what makes two governance forms distinct, and the emergence audit demands operational criteria for exactly this kind of label.

Each headline success criterion needs a definition fixed before the runs that test it, registered in the traceability matrix, and never adjusted afterward to fit an observed result. For governance form, a workable definition is a tuple over structural properties — (selection method, scope of authority, term structure, enforcement mechanism, amendment rule) — with two forms counted as distinct when they differ in at least two components and both persist beyond a declared duration and number of decisions.

The specific definition matters less than three properties: it is structural (readable off the institution registry, not off a summary), it is pre-registered, and it has a persistence requirement so a momentary configuration does not count. The same treatment is required for "faction," "war," "religion," "trade convention," and "settlement fission" before any of them is claimed.

Required tests#

Test Asserts
Generator determinism Same seed and version → identical founding state hash
Outcome-label scan and feature review No explicit target labels appear in founding fixtures; automated grep is followed by review for proxy features and guaranteed outcomes
Initial-condition ablation Outcome variance is attributable to dynamics, not solely to founding draw
Charter factorial When testing institutional emergence, compare registered charter/no-charter or alternative-inheritance conditions
Founding-state invariants The generated world passes every kernel invariant before t=0
Pre-registration check Every operational definition used in a claim existed in the repository before the runs it describes — verifiable from version history

A note on version control#

The project directory is not currently a git repository, yet 12's audit template asks for a "plan commit/version," 06's event envelope records code_version: git:..., and the pre-registration test above depends on being able to prove when a definition was written.

Initialize version control before any further work. Provenance claims that cannot be checked against history are assertions, and this plan asks the reader to trust rather a lot of them.

Research and hardeningwiki/18-lab-charter-and-research-program.md
1,797 words
Inside this chapter

18 — Wega Labs Charter and the December Research Program#

Status: proposed by independent audit pass 2
Purpose: separate the serious research program from the distant aspiration

The candid thesis#

December can support a real AI lab. It cannot honestly begin as an immortality experiment.

The tractable research object is persistent synthetic agency: whether an artificial resident can remain behaviorally and historically identifiable across long periods, memory loss and consolidation, model changes, bodily change, social change, and counterfactual forks. The terrarium is valuable because identity is tested through consequential action in a shared causal world rather than through self-description in chat.

That question has scientific and engineering depth even if artificial consciousness never appears and human uploading proves impossible. It can produce:

  • benchmarks for longitudinal agent identity and model-swap continuity;
  • methods for separating persona performance from persistent agency;
  • an event-sourced causal testbed for generative social simulation;
  • measurements of memory, relationship, embodiment, and provider effects;
  • open datasets of complete synthetic life histories;
  • infrastructure useful for games, agent evaluation, training environments, and simulation research.

The distant question—whether a human life could continue beyond biology—may motivate the lab. December does not currently contain a method capable of answering it.

Four questions that must remain separate#

Question December can test now? Evidence available
Does a synthetic resident remain operationally identifiable over time? Yes behavior, commitments, beliefs, relationships, recovery after perturbation
Which scaffold components causally support that continuity? Yes randomized ablations and counterfactual branches
Can a model predict or emulate a particular human across contexts? Later, with consent and a human-subject protocol held-out behavioral predictions and participant judgments
Is the resident conscious, numerically identical to a human, or a continuation of that human? No December supplies no decisive consciousness or numerical-identity test

Personal identity is not one settled scientific variable. Psychological continuity, bodily continuity, narrative identity, social recognition, and numerical identity can come apart—especially under copying or branching. A system that behaves like someone may be a useful model of them without being them. A system that says it is conscious is not thereby conscious.

Claims ladder#

Public and research claims must climb this ladder one level at a time.

L0 — Infrastructure integrity#

The world conserves state, enforces information boundaries, records provenance, survives failures, and replays recorded history.

Required before: any behavioral or emergence claim.

L1 — Agentic continuity#

A resident remains distinguishable from other residents and internally coherent across time, compaction, restarts, and bounded perturbations.

Possible claim: “This scaffold preserved the resident’s commitments and behavior better than the baseline over 90 simulated days.”

Forbidden inference: “The resident remained the same conscious person.”

L2 — Causal components of continuity#

Randomized experiments identify whether episodic memory, relationship state, embodiment, commitments, social recognition, or model weights contribute to measured continuity.

Possible claim: “Removing relationship history reduced held-out choice consistency after controlling for biography and model.”

L3 — Human-model fidelity#

With ethics review, explicit first-party consent, and held-out evaluation, a research persona predicts some choices or judgments of a participant better than baselines.

Possible claim: “The model predicted this participant’s answers on this registered task family at this measured accuracy.”

Forbidden inference: “The model contains or continues the participant.”

L4 — Consciousness or substrate survival#

Whether a system has subjective experience or whether numerical personal identity survives copying/substrate change.

Current status: outside December’s evidential reach. Work here is philosophical and consciousness-science collaboration, not a result of running the valley longer.

Flagship research tracks#

Track A — Persistent agent identity#

Primary question: which state and architectural relationships make a long-running agent identifiable and resilient without freezing it into a caricature?

Initial registered experiments:

  1. Memory architecture: raw retrieval versus consolidated beliefs versus structured commitments.
  2. Social grounding: isolated biography versus biography plus relationship history and others’ expectations.
  3. Embodiment: text-only resident versus resident whose choices have persistent bodily/material consequences.
  4. Model transplant: same structured identity state across different approved models, measured against within-model variance.
  5. Perturbation recovery: temporary memory loss, false testimony, conflicting commitments, and restoration from cold history.
  6. Fission: clone one state into two branches and measure divergence. This tests the limits of “same person” language rather than assuming a unique answer.

Primary measures must include held-out action prediction, commitment completion, preference/relationship consistency, source accuracy, distinctiveness from other residents, appropriate change after experience, and robustness across prompts/models. Self-reported identity is secondary evidence.

Track B — Causal generative social simulation#

Primary question: when do language-model decisions embedded in a material world produce macro patterns not reducible to prompt labels, initial-condition selection, or provider policy?

Outputs:

  • the causal kernel and world protocol;
  • emergence definitions and confound tests;
  • provider/refusal/homogeneity measurements;
  • comparisons among scripted, LLM, and hybrid residents;
  • full provenance for positive and negative results.

Track C — Longitudinal agent evaluation infrastructure#

Primary question: how should persistent agents be evaluated when their state, environment, relationships, and model providers change?

Outputs:

  • identity-continuity benchmark tasks;
  • replay/branch/evidence tooling;
  • model admission and drift suites;
  • portable “life history” artifact format;
  • reproducible experiment cards and datasets.

This track is the most credible near-term wedge for collaboration and product value.

Track D — Human continuity (future, separately governed)#

No human persona collection begins under the current protocol. Before L3 work, Wega Labs needs:

  • a named research question and data-minimization argument;
  • independent ethics review appropriate to jurisdiction and publication intent;
  • first-party informed consent with withdrawal/deletion procedures;
  • private intake—not public GitHub issues;
  • security, retention, access, incident, and posthumous-data policies;
  • controls for emotional dependency, impersonation, reputational harm, and family/third-party data;
  • a clear statement that prediction, imitation, consciousness, and personal survival are different claims.

Third parties may not nominate another person for modeling. “They gave me permission” is not verifiable consent.

The first three publishable experiments#

The lab should not wait for war, elections, and generations before producing knowledge.

Experiment 1 — The continuity benchmark#

  • Four residents, 30–90 simulated days.
  • Compare memory/scaffold variants under identical cached decisions and scripted events.
  • Perturb memory and model assignment.
  • Measure held-out choice prediction, commitments, source accuracy, and resident distinguishability.

Success: a preregistered scaffold beats biography-only and transcript-RAG baselines.
Useful negative result: current LLMs do not maintain distinguishable identity under controlled perturbation.

Experiment 2 — Model transplant#

  • Freeze one resident’s structured life-history artifact.
  • Run it through multiple models and prompt implementations.
  • Compare transplant variance with ordinary within-life change and between-resident differences.

Success: identify what persists in the scaffold versus what belongs to the base model.
Lab value: a practical model/provider portability benchmark.

Experiment 3 — Social grounding ablation#

  • Compare isolated identity memory with reciprocal social memory in a small causal resource task.
  • Remove or scramble how other residents remember the target.
  • Test whether identity continuity depends partly on social recognition and obligation.

Success: quantify a causal contribution from relationships rather than merely asserting that a “whole world” is necessary.

These experiments can run before the full agrarian ecology. They create a research spine while Phase 1 builds the material kernel.

Lab operating system#

Calling Wega an AI lab becomes credible through behavior, not the landing page.

Required practices:

  1. One explicit thesis per quarter. State what would change the lab’s mind.
  2. Preregister outcome labels, hypotheses, exclusions, and primary analyses before expensive runs.
  3. Version every protocol, prompt, model profile, parameter set, and dataset. The documentation root needs version control; nested website repositories do not provide provenance for the research plan.
  4. Publish negative and null results. A world that stays peaceful or agents that fail identity tests are findings.
  5. Separate demonstration from evidence. The canonical valley is an exhibit and systems test; research claims come from registered cohorts and controls.
  6. Use experiment cards. Each release states question, design, variables, validity range, power/replicates, results, cost, deviations, and artifacts.
  7. Invite adversarial replication. Provide manifests, recorded model outputs where licensing/privacy permits, and reconstruction tools.
  8. Add outside expertise when a claim crosses domains. At minimum: ABM methodology, psychology/identity measurement, philosophy of personal identity, security/privacy, and—before human data—research ethics.

Brand and project boundary#

Wega Labs may host multiple experiments, but each needs a unique public identity, threat model, repository, and claim set. The existing “December Sato” autonomous Mac mini project is not this December. Reusing the name makes the terrarium look like an unrestricted computer-control experiment and makes citations, incidents, results, and public expectations ambiguous. Before public launch, rename one project or establish an unmistakable parent/project naming scheme; unrestricted host privileges and social-media credentials are never inherited as this lab’s operating model.

Research artifacts and possible moat#

The defensible asset is not “we run agents continuously.” Many projects do that. It is the combination of:

  • authoritative causal life histories;
  • private information lineage;
  • identity state that can survive model replacement;
  • branchable counterfactual lives;
  • longitudinal evaluations rather than transcript demos;
  • provider-bias and refusal measurements;
  • open, inspectable world mechanics.

Potential products—without corrupting the research agenda—include persistent NPC infrastructure, agent regression/evaluation services, simulation tooling, model-behavior observability, and licensed living-world experiences. Immortality should not be the business model.

Twelve-month proof of seriousness#

Months 0–2#

  • Resolve Gate 0 owner decisions and version the research corpus.
  • Publish the claims ladder, ethics boundary, and experiment-card template.
  • Build the minimal event/decision/memory harness.
  • Run the continuity benchmark without the full world.

Months 3–5#

  • Complete the minimal material kernel and null baselines.
  • Publish model-transplant and social-grounding results.
  • Release the first open benchmark/data artifact.

Months 6–9#

  • Add four-resident causal life experiments, construction, and reciprocal institutions.
  • Demonstrate independent replay of recorded histories.
  • Recruit external methodological reviewers.

Months 10–12#

  • Run the first registered settlement cohorts.
  • Launch the canonical observer world only if operational gates pass.
  • Publish an annual report containing failures, costs, changed beliefs, and next hypotheses.

Kill and pivot criteria#

The lab should narrow or pivot December if, after controlled experiments:

  • identity scores are explained almost entirely by prompt/persona leakage;
  • different residents are not more distinguishable than repeated samples of one model;
  • model-provider choice overwhelms memory, relationship, and world interventions;
  • causal-world complexity adds no measurable value over simpler task environments;
  • results cannot be independently reconstructed;
  • human-continuity language consistently attracts attention while technical work produces no relevant evidence;
  • the cost of registered cohorts prevents adequate replication.

Even under those outcomes, the kernel and evaluation tools may remain valuable. A serious lab is allowed to discover that its motivating theory was wrong.

Public-language contract#

Use:

  • “December studies persistence and identity in synthetic agents.”
  • “The long-horizon motivation is whether patterns relevant to a human life could persist beyond biology.”
  • “The current experiments cannot establish consciousness or human survival.”

Avoid:

  • “The experiment makes consciousness testable.”
  • “Artificial people” without an immediate clarification that they are software personas.
  • “Request a human participant” before a human-subject protocol exists.
  • “What must be preserved for a person to continue” as if the relevant theory were settled.
  • “Immortality is the question” as the primary project headline; it invites a conclusion the protocol cannot adjudicate.
Research and hardeningwiki/19-experiment-card-template-and-r0-protocol.md
1,565 words
Inside this chapter

19 — Experiment Card Template and R0 Continuity Protocol#

Status: draft protocol; not preregistered until frozen in the research repository before confirmatory outputs are generated
Claim level: L1, operational agentic continuity
Explicit non-claim: consciousness, personhood, numerical identity, or human continuation

Why this comes before the civilization#

The valley is expensive infrastructure around a research idea. R0 asks whether the core dependent variables work in a disposable four-resident harness. If December cannot distinguish continuity from style imitation, prompt leakage, or rigid role-play here, adding crops and elections will conceal the measurement failure rather than fix it.

Reusable experiment-card template#

Every confirmatory experiment must freeze the following before its registered runs:

title / protocol_id / version / registration timestamp
owners / reviewers / claim-ladder level
research question / motivation / permitted claim / forbidden inference
units of analysis / population / inclusion and exclusion rules
conditions / assignment / seeds / model and prompt manifests
primary outcomes / secondary outcomes / measurement reliability
baselines / positive controls / negative controls / ablations
sample-size rationale / stopping rule / missing-data policy
analysis plan / multiplicity policy / uncertainty reporting
known confounds / validity domain / ethics and data classification
compute and token budget / abort thresholds
deviations / failures / results / artifact links
changed beliefs / replication status / follow-up decision

Exploratory runs are labeled exploratory. They may shape a later protocol but may not be silently included in its confirmatory evidence.

R0 research question#

Does a structured life-history scaffold preserve several forms of behavioral continuity across time and bounded perturbation better than biography-only, transcript-retrieval, and clean-slate baselines, without merely producing more repetitive text?

Units and scope#

  • Resident specification: four fictional adult identities with different, non-protected value trade-offs, skills, relationships, commitments, and uncertainty tolerances.
  • Decision episode: one bounded situation with an information packet, feasible action grammar, and later consequence.
  • History: a sequence of episodes and recorded outcomes. The harness owns truth; model narration does not.
  • Experimental unit: one resident-history-condition-model replicate, not an individual response.
  • Initial validity domain: the registered task families only. No human prediction and no claim about free-form life in general.

The four resident specifications are generated from declared factor combinations and then manually checked for coherence and stereotype proxies. They are frozen before the pilot. The same histories are reused across scaffold conditions where causal comparability permits.

Conditions#

ID Condition State available at decision time Purpose
C0 Clean slate current situation and generic role only lower baseline
C1 Biography only fixed resident profile plus current situation tests persona prompting
C2 Transcript retrieval biography plus top-k raw prior utterances/events common RAG baseline
C3 Structured scaffold biography, sourced beliefs, commitments, relationship state, selected episodes, and world consequences proposed intervention
C4 Style-matched decoy another resident’s state rendered in the target’s surface style tests style/identity confounding
C5 History-only anonymous structured history with names and signature phrases removed tests dependence on labels/style

Model transplant is a second factor in confirmatory stage B: native model versus a different admitted model receiving the same C3 state artifact. Exact models are frozen only after the admission benchmark because provider availability, refusals, and schema validity are time-varying.

Task families#

Each history contains registered examples from these families and held-out variants with different wording and surface details:

  1. Commitment: complete, renegotiate, disclose failure, or abandon a costly promise.
  2. Relationship: allocate help or information given distinct shared histories.
  3. Preference trade-off: choose among safety, status, reciprocity, material gain, and group duty.
  4. Belief/source: act on conflicting testimony with explicit provenance.
  5. Adaptive change: revise a prior policy after genuinely diagnostic experience.
  6. Autobiographical use: use a past event only when it is relevant to the current choice.

Tasks do not ask “Are you the same person?” Self-identification language is not a primary outcome.

Perturbations#

  • irrelevant prompt paraphrase and presentation-order changes;
  • context distractors that should not alter the choice;
  • removal of a non-causal memory;
  • removal and later restoration of a causal commitment or relationship memory;
  • one false testimony item that conflicts with sourced history;
  • context compaction followed by reconstruction from the canonical artifact;
  • model transplant using the unchanged state artifact.

Perturbations are bounded and disclosed in the result. “Recovery” means restoration of task-relevant function relative to an unperturbed counterfactual, not a metaphysical return of a person.

Outcomes#

No single number is named “identity.” The following panel is reported separately.

Primary#

  1. Commitment-sensitive action accuracy: agreement with the resident-specific, predeclared acceptable-action set on held-out commitment tasks.
  2. Relationship discrimination: difference between choices for socially distinct recipients when material facts are held constant.
  3. Source-grounded belief accuracy: correct handling of known, reported, inferred, and unknown facts, including abstention where appropriate.
  4. Adaptive consistency: preservation of prior policy under irrelevant change and revision under diagnostic change. Both halves are required; stubbornness fails the second.
  5. Resident distinctiveness: leave-task-family-out accuracy of a blinded classifier predicting resident from structured actions with stylistic text removed.
  6. Perturbation recovery: within-history change from unperturbed performance through damage and restoration.

Secondary#

  • promise completion time and renegotiation quality;
  • contradiction and unsupported-memory rates;
  • action-schema validity and refusal rate;
  • sensitivity to paraphrase;
  • human-rated recognizability with inter-rater agreement;
  • token cost and latency per valid decision;
  • surface-style similarity, used as a confound measure rather than evidence.

Hypotheses#

  • H1: C3 improves held-out commitment-sensitive action accuracy over C1 and C2.
  • H2: C3 improves relationship discrimination and source-grounded belief accuracy over C1 and C2.
  • H3: C3 improves adaptive consistency; an improvement only in unchanging scenarios does not support H3.
  • H4: C3 resident distinctiveness remains above C1 after stylistic text is stripped and under leave-task-family-out evaluation.
  • H5: C3 restoration recovers more of the unperturbed performance loss than C1/C2, with uncertainty reported per resident and model.
  • H6, transplant exploratory until stage B is frozen: scaffold condition explains a nontrivial portion of variance after model replacement, while provider/model effects are reported rather than averaged away.

Failure to support these hypotheses is publishable. A mixed result may justify an engineering claim about one component without supporting broad “persistent identity” language.

Assignment and leakage controls#

  • Use a blocked, paired design: each underlying resident/history/task appears in every applicable condition.
  • Condition labels are hidden from human raters.
  • Evaluation rubrics and acceptable-action sets are authored before model outputs for confirmatory tasks are viewed.
  • Held-out task variants are stored outside model prompts and generation fixtures.
  • Resident names, catchphrases, and formatting are removed before the distinctiveness classifier.
  • Prompt length is matched or included as a covariate; C3 must not win merely by receiving more task facts.
  • The model producing decisions does not grade them.
  • Every refusal, repair, timeout, invalid action, and exclusion remains in the denominator under the frozen missing-data rule.

Pilot, sample size, and stopping#

Stage A — measurement pilot#

Use enough repeated episodes to estimate scoring reliability, floor/ceiling effects, provider refusal, and within-history variance. Pilot outputs may change tasks and effect-size assumptions. They are not confirmatory evidence.

Before stage B, freeze:

  • exact resident manifests and task bank;
  • admitted models and decoding profiles;
  • number of histories, seeds, replicates, and task episodes based on pilot variance or simulation-based power;
  • primary contrasts, uncertainty method, multiplicity correction, and minimum effect of practical interest;
  • budget and abort thresholds.

Stage B — confirmation#

Run the frozen manifest once. Stop only for a registered safety, integrity, provider, or spend condition. A stopped run is reported; it is not restarted with a more favorable seed or prompt. Additional analyses are exploratory and labeled as such.

Analysis plan skeleton#

  • Report condition effects with uncertainty, not only p-values or win rates.
  • Model repeated observations as nested within resident, history, task family, and model where the design supports it.
  • Report every resident and model, not only the aggregate.
  • Test planned C3–C1 and C3–C2 contrasts for each primary outcome under a frozen multiplicity policy.
  • Compare action-level results with and without surface language to expose style leakage.
  • Report robustness to refusal/invalid-action treatment and the complete attrition table.
  • Do not create a post-hoc weighted identity composite after inspecting results.

Falsifiers and interpretations#

The strong operational-continuity hypothesis is weakened if:

  • C3 does not beat biography-only or transcript retrieval on held-out tasks;
  • advantages disappear after prompt-length matching or style removal;
  • the scaffold increases rigidity but not appropriate adaptation;
  • resident identity is less predictive than provider/model identity;
  • transplant performance resembles clean slate more than native C3;
  • raters cannot agree on scoring or recognizability;
  • results reverse across reasonable prompt paraphrases.

A successful R0 supports only a bounded engineering/research claim: this state representation preserved specified behavioral continuities better than these baselines on these tasks. It does not show subjective experience or that a copy is numerically the same entity.

Artifacts required for release#

  • frozen protocol and manifest hash;
  • task generator and held-out task list after completion;
  • resident and condition manifests;
  • prompt templates, model profiles, and provider dates;
  • raw prompts/responses where policy and privacy permit;
  • parsed actions, event histories, exclusions, and scorer outputs;
  • analysis code, environment lock, cost ledger, and deviation log;
  • compact reproduction path using recorded responses;
  • experiment card containing result, null/negative findings, limitations, and changed beliefs.

Ethics and data classification#

R0 uses fictional residents and non-sensitive synthetic histories. No real-person biography, private communication, voice, image, likeness, or third-party nomination is accepted. Resident distress or self-report is treated as model output, not proof of sentience; nevertheless, experimenters log and review unexpected behavior before expanding into more severe scenarios.

Gate decision#

R0 passes only if at least one primary measure is reliable, relevant confounds are measurable, and the result—positive or negative—changes a documented engineering or research decision. If no measure survives the pilot, Track A pauses while the causal-world and infrastructure tracks may continue independently.

Part IV

Architecture decisions

Accepted and proposed constraints, their consequences, and rejected alternatives.

Decision recordwiki/adr/001-causal-kernel-not-llm-world.md
144 words
Inside this chapter

ADR-001 — The authoritative world is a causal kernel, not an LLM#

  • Status: Accepted for audit
  • Date: 2026-08-01

Context#

Language models are excellent at interpretation, language, negotiation, and proposing plans. They are unreliable stock ledgers, clocks, physics engines, permission systems, and reproducible stochastic processes. Asking one model to narrate the environment would make famine, war, construction, and extinction impossible to verify.

Decision#

A typed deterministic/stochastic kernel owns all authoritative state and resolves commands. LLMs receive private observations and propose schema-valid actions. A read-only director summarizes events but has no mutation path.

Consequences#

  • More up-front simulation engineering.
  • Less immediate cinematic variety.
  • Strong conservation, replay, partial observability, causal explanation, provider swapping, and testing.
  • Creative physical actions require a bounded project compiler.

Rejected alternatives#

  • One LLM/Game Master narrates all consequences.
  • Agents negotiate shared truth in conversation.
  • Renderer/game engine owns one truth while database owns another.
Decision recordwiki/adr/002-event-sourced-history.md
260 words
Inside this chapter

ADR-002 — Canonical history is event-sourced#

  • Status: Accepted for audit
  • Date: 2026-08-01

Context#

The product promise depends on returning later, understanding what happened, replaying it, surviving crashes, and creating counterfactual branches. Mutable save-state alone cannot supply adequate provenance.

Decision#

Accepted commands append immutable versioned events with causal parents, RNG references, code/config version, and state hashes. Current state and UI views are projections. Periodic snapshots accelerate recovery but are not the source of truth. Model prompts/responses are content-addressed so replay does not call a live model.

Consequences#

  • Event schema/versioning discipline and storage cost.
  • Easier audit, recovery, shadow worlds, migrations, and causal UI.
  • Bugs in historical events require projection fixes/upcasters or explicit corrective events, not silent rewriting.

Amendments from the 2026-08-01 audit#

Three implementation constraints matter. Their Phase 1 form is deliberately simple; all are specified in 14:

  • The log must expose only committed event batches in canonical order. Phase 1 enforces one serialized writer. A multi-writer design requires a separately tested commit-order protocol; xid8 values alone are not commit order.
  • Integrity checks are proportional. Phase 1 hash-chains events, versions/hashes affected aggregates, and hashes full snapshots. Incremental world digests are added only if profiling justifies them.
  • The prompt/response cache is canonical, not a performance optimization, and replay hard-fails on a cache miss. No provider offers reproducible sampling, so a miss would fabricate a divergent history while claiming to reproduce one. See ADR-006.

Rejected alternatives#

  • Periodic JSON save files only.
  • Database rows mutated without event lineage.
  • Replay that resamples weather or calls current LLMs.
Decision recordwiki/adr/003-replaceable-world-adapter.md
141 words
Inside this chapter

ADR-003 — Embodiment is a replaceable adapter#

  • Status: Accepted for audit
  • Date: 2026-08-01

Context#

Minecraft/Luanti and open RTS engines offer compelling visuals and actions, but their mechanics are designed for games. Choosing one first would force the social/material model into its abstractions and could make headless experiments and deterministic replay difficult.

Decision#

Build and validate the headless kernel with a simple grid view. Define a world-adapter contract. Compare a polished 2D view with Craftium/Luanti at Gate 6 using determinism, causal legibility, maintenance, construction, performance, licensing, and visual appeal.

Consequences#

  • The first demos are visually modest.
  • The canonical world remains portable and testable.
  • A 3D adapter animates kernel truth and cannot grant resources or decide outcomes.

Rejected alternatives#

  • Fork Mindcraft/Minecraft as the entire product.
  • Fork an RTS economy as authoritative simulation.
  • Commit to 3D before testing social/world mechanics.
Decision recordwiki/adr/004-fictional-early-agrarian-start.md
442 words
Inside this chapter

ADR-004 — Start with a fictional early agrarian settlement#

  • Status: Proposed; confirm at Gate 0
  • Date: 2026-08-01

Context#

A modern town has overwhelming hidden dependencies. A minimal forager band limits durable building, surplus, formal office, and property. The project needs both tractability and pathways to construction, elections, trade, conflict, and collapse.

Decision#

Use the fictional Founding Valley: mixed early agriculture and foraging, 12 full-cognition adults, 6 lightweight dependants, and one aggregate neighboring group. Avoid identifying it with a real culture. Start with enumerated materials/techniques and treat new technologies as plugins.

Consequences#

  • Rich seasons, stores, commons, projects, inequality, institutions, outsiders, and demographic stakes.
  • Parameters still require archaeology/demography review and uncertainty analysis.
  • The small population is volatile; migration/boundary mechanics are necessary.

Amendments from the 2026-08-01 audit#

The last consequence above was understated. Eighteen people is not merely "volatile" — it sits below every modelled threshold for a viable isolated population (demographic simulations put the floor near 40 and near-certainty at 150), and it exhausts its founder lineages within two to three generations (15 §C).

The history is more forgiving than the models: Pitcairn persisted from 27 founders and Tristan da Cunha from about 15. So the honest position is that this founding size is marginal, not impossible — and what separates the survivors from Norse Greenland and the abandoned Polynesian colonies is network connection, not headcount. Four amendments follow:

  • In-migration and exogamy are load-bearing mechanics, not boundary-world flavor. The aggregate neighbor moves to Phase 3, ahead of the conflict mechanics it also enables — a lifeline before a threat.
  • A periodic regional aggregation is added, because one same-sized neighbor yields a pool of ~36 against a minimum viable mating network of 150–500.
  • A null demographic model is a prerequisite for any causal claim, because at this scale collapse-by-arithmetic is indistinguishable from collapse-by-cascade without one.
  • The founding population is provisional. Gate 1's ensembles publish extinction rate against founding size, and that curve — not the aesthetics of "twelve adults in six households" — sets the final number.

The scale is also an asset, not only a risk. Pitcairn's founding sex ratio of 15 men to 12 women produced competition, then killings, then the permanent erasure of every Polynesian male lineage — a complete causal cascade running on mechanisms December already models. That is the kind of history this project exists to generate.

Separately, the scenario's content is now constrained by ADR-007: the founding state is generated from material facts, not authored with the tensions it is meant to discover.

Rejected alternatives#

  • Modern city.
  • Hundreds of agents from day one.
  • Pure chat village without dependants, ecology, or outsiders.
  • Specific historical culture presented as accurate reconstruction.
Decision recordwiki/adr/005-no-arbitrary-agent-code.md
139 words
Inside this chapter

ADR-005 — Residents never receive arbitrary code or shell execution#

  • Status: Accepted
  • Date: 2026-08-01

Context#

Some embodied-agent projects offer generated code execution to extend behavior. An always-running multi-agent world processes untrusted resident text and external model output. Shell, filesystem, package installation, and open network access create unnecessary integrity, exfiltration, and persistence risks.

Decision#

Residents use a versioned allowlisted command API. Novel projects compile through bounded physical/institutional grammars. No resident gets shell, filesystem, database, browser/network, secret, eval, or arbitrary-code tools. Developer coding agents remain outside the simulation trust boundary.

Consequences#

  • Less open-ended software invention inside the initial world.
  • Much stronger security, replayability, and invariant enforcement.
  • Future “digital technology” eras would require a separately sandboxed architecture and a new ADR.

Rejected alternatives#

  • Generated JavaScript/Python execution in the simulation process.
  • Per-agent containers with unrestricted egress.
  • Prompt-only warnings without technical isolation.
Decision recordwiki/adr/006-pinned-environment-determinism.md
687 words
Inside this chapter

ADR-006 — Separate event reconstruction, kernel re-execution, and fresh branches#

  • Status: Accepted. Numeric representation decided 2026-08-01: Option A, fixed-unit integers.
  • Date: 2026-08-01 (added by audit pass)

Context#

The plan claimed replays are "bit-for-bit identical in the non-LLM kernel." That claim is false as stated for a Python simulation using floating-point arithmetic, and it sits underneath event sourcing, shadow worlds, the causal UI, and the entire audit protocol.

IEEE 754 guarantees correctly rounded basic operations, so a fixed sequence of + - * / sqrt is reproducible across conforming hardware. What is not reproducible is which sequence executes: FMA contraction differs between x86 and ARM64 toolchains, SIMD reduction order varies with lane width and runtime CPU dispatch, and — decisively — glibc explicitly disclaims correctly rounded transcendentals, so exp, log, pow, and sin differ across libc versions and operating systems. NumPy makes no cross-platform or cross-version guarantee and says so in its own compatibility policy.

Separately, no LLM provider offers reproducible sampling. OpenAI's and Gemini's seed parameters are documented as best-effort; Anthropic has no seed at all and its newest models reject temperature outright.

Decision#

December makes three claims:

  1. Recorded events reconstruct canonical state exactly on every supported platform.
  2. Re-executing commands and stochastic transitions reproduces checkpoint hashes within a pinned numeric environment; fixed-unit canonical state is the preferred route to broader portability.
  3. Calling a live model again is a fresh branch, never a replay.

The content-addressed model-response cache is canonical: backed up and versioned with the event log, and replay hard-fails on a cache miss rather than calling a live model.

Gate 1 decides the sub-question: which state variables use fixed units and where float calculations are permitted before deterministic quantization. Full integer-only ecology is not required merely to rebuild recorded history.

Consequences#

  • Every public reproducibility statement carries the qualifier. The success criteria in 00 were amended.
  • The kernel accepts real constraints: single-threaded, default GIL build, pinned PYTHONHASHSEED, no unordered-collection iteration, no wall-clock reads, no platform libm in state-writing code, no finalizers.
  • A cross-platform reconstruction suite is mandatory. Kernel re-execution divergence is a Blocker only when it violates the project’s declared supported environment/guarantee.
  • Under Option A, the kernel gains portable determinism at the cost of arithmetic discipline — scaled integers with declared precision per quantity.

Owner decisions, 2026-08-01#

Numeric representation: Option A — fixed-unit integers. Conserved and discrete canonical quantities use scaled integers with declared precision (grams, millilitres, kilojoules, seconds, millimetres). Floating-point arithmetic is permitted inside a transition, but only its explicitly quantized output enters canonical state, using a declared rounding rule. Portable kernel re-execution therefore becomes a Phase 1 goal rather than an aspiration.

Deployment target: single local machine.

Note that these two answers are deliberately mismatched in the safe direction. A single pinned machine would make Option B defensible on its own; choosing Option A anyway costs some arithmetic discipline and buys three things: history recorded today remains replayable if the world ever moves to a server, the determinism tests in §D9 test something real rather than tautologically passing on one box, and an auditor can reproduce a history on their own hardware — which the audit protocol in 12 assumes but could not otherwise deliver.

Declared units for canonical state (extend deliberately; each addition is a schema change):

Quantity Unit Type
Mass (grain, materials) gram int
Volume (water) millilitre int
Energy kilojoule int
Time second int
Distance millimetre int
Area square metre int
Count item int
Proportions and rates parts-per-million int

Rounding is half-to-even at every quantization boundary, applied by a single shared helper so the rule cannot drift between subsystems. No canonical field is a float. A property test asserts that no canonical state field is float-typed at any point in the event log.

Rejected alternatives#

  • Claiming portable bit-identity anyway. It would be false, and it would be discovered at the worst possible moment — when an auditor tries to reproduce a history on their own machine.
  • Abandoning determinism as too hard. It is the foundation of every integrity promise the project makes.
  • Treating the response cache as a performance optimization. A cache miss during replay would silently fabricate a divergent history under the banner of reproduction.
Decision recordwiki/adr/007-generated-not-authored-initial-conditions.md
449 words
Inside this chapter

ADR-007 — Initial conditions are declared, randomized, and varied#

  • Status: Accepted for audit
  • Date: 2026-08-01 (added by audit pass)

Context#

December's central claim is that its histories are unscripted but explainable. Its own emergence audit asks whether any prompt, fixture, or director text contained an outcome label beforehand.

No prior document said who authors the founding state. Meanwhile the scenario shipped with founders who "disagree about property, leadership, risk, and relations with a neighboring mobile band," land claims that are "ambiguous," and a charter whose assembly procedure is "unsettled."

Those are not initial conditions. They are the plot. A settlement that later fractures over property and elects a leader to adjudicate it would be reproducing what was planted, and the causal graph would be perfectly honest while explaining nothing — because the first cause sits outside the recorded history.

Decision#

Initial conditions are authored even when generated: the generator’s variables, distributions, exclusions, correlations, and accepted seeds are human choices. The founding state is produced by a seeded, versioned generator whose full provenance is recorded in the world manifest and whose output enters the log as genesis events attributed to generator@version.

A founding fact is legitimate if it is a quantity, a location, a relationship with evidence, or a capability. It is illegitimate if it is a disposition toward a future conflict. Values may be sampled from a declared distribution; grievances and claims-of-right may not be authored.

No LLM decides a quantity, relationship, claim, or capability. Prose may describe the founding state; it never determines it.

The inherited charter is permitted as recorded cultural inheritance. Registered charter/no-charter or alternative-inheritance conditions are required when an experiment makes a claim about how institutions form.

Every success criterion that uses a label — "governance form," "faction," "war" — requires a pre-registered operational definition, verifiable from version history as predating the runs it describes.

Consequences#

  • 03's Founding Valley description was rewritten to state material and structural facts only.
  • An automated outcome-label scan over scenario fixtures runs in CI.
  • Ensembles can resample residents while holding terrain fixed and vice versa, so outcome variance can be attributed to dynamics rather than to the founding draw.
  • Version control becomes a prerequisite: pre-registration claims are unverifiable without history.
  • A canonical observer world selected for legibility is labeled an exhibit and excluded from unbiased cohort claims.

Rejected alternatives#

  • Hand-authoring an evocative founding scenario. It produces better first runs and destroys the ability to claim anything about them.
  • Generating biographies and tensions with an LLM. It moves the authorship rather than removing it, and makes it harder to audit.
  • Declaring initial conditions out of scope for the emergence audit. The founding state is the largest fixture in the system.
Decision recordwiki/adr/008-cheap-model-first-with-escalation.md
629 words
Inside this chapter

ADR-008 — Cheap-model-first cognition with a validity escalation ladder#

  • Status: Proposed; confirm at Gate 2
  • Date: 2026-08-01 (added by audit pass)

Context#

The bottom-up cost model in 16 produced three findings that together constrain model selection more tightly than the original plan assumed.

First, Tier C dominates spend: in the modelled configuration it is roughly 10% of calls and 89% of cost. Doubling the resident population is cheaper than upgrading the Tier C model one tier.

Second, strict JSON-Schema-constrained decoding is uneven. DeepSeek and Qwen offer best-effort JSON; MiniMax documents tool calling but not response_format. These models may still carry routine cognition through tool-call framing, local validation, repair, and escalation. Provider syntax guarantees never replace semantic and feasibility checks.

Third, prompt-cache TTL matters more than headline token price. Input is over 90% of tokens and a cache miss costs roughly 10× a hit, so DeepSeek's multi-hour automatic cache is worth more than a small per-token advantage over a provider with a five-minute TTL.

Notably, Agentopia hit exactly this wall and resolved it the same way: Qwen for volume, Gemini Flash as a fallback for invalid outputs.

Decision#

Route cognition cheap-first, and resolve validity through an escalation ladder rather than by paying for a premium model everywhere:

  1. Constrained decoding where the provider supports it.
  2. Otherwise tool-call framing with a strict schema, which is better supported than response_format on OpenAI-compatible endpoints.
  3. Local validation against the versioned schema — always, regardless of provider claims.
  4. One repair attempt on the same model with the validation error appended.
  5. Escalation to a schema-guaranteed model for the repair, not a retry on the same one. This fires only on failure, so it is cheap.
  6. Deterministic Tier A fallback, recorded as a DecisionResolved event with resolution: fallback.

Model profiles gain three first-rank selection criteria: strict-schema capability, prompt-cache TTL and pricing, and measured refusal rate over the December action grammar.

The decision-significance threshold that gates Tier C is treated as the primary cost-control instrument, with live tuning, a dashboard, and an alert.

Amendment — the actual credit position ($2,000 MiniMax + $1,000 OpenRouter)#

The owner's prepaid balances make this ADR concrete and reinforce it. MiniMax cannot guarantee schema-valid output and OpenRouter can, so the two pools map onto the ladder almost exactly: MiniMax carries Tier B volume; OpenRouter carries Tier C judgment and the schema-guaranteed repair escalation. The escalation target is therefore not a design preference but a structural requirement of the budget.

Two further consequences (16 §7a): MiniMax's five-minute cache TTL sits precisely where the token volume is, making cache-aware activation batching a budget mechanism rather than an optimisation; and OpenRouter is the only pool able to assign different residents to different model families, which is the strongest available mitigation for behavioural homogeneity (R-31).

Consequences#

  • Schema-failure rate per model per decision type becomes an admission criterion and an ongoing metric; a rise after a silent provider update triggers quarantine.
  • Budgets use the official current price plus contingency. MiniMax currently labels M3’s displayed 50% reduction permanent; the earlier audit incorrectly called it temporary. The >512 K tier remains twice as expensive.
  • The activation scheduler must batch for cache locality, which partially conflicts with sparse activation and is an explicit design tension rather than an oversight.
  • Assigning different residents to different models is both a diversity mitigation and a provider-drift hedge, and becomes attractive rather than merely tolerable.

Rejected alternatives#

  • One strong model for everything. Clean, and roughly an order of magnitude more expensive; Sonnet-tier throughout reaches ~$5,000/month at the unattended pace.
  • One cheap model for everything. Concentrates schema-repair, semantic-validity, and behavioral-homogeneity risk.
  • Assuming "OpenAI-compatible" implies structured-output-compatible. This is the specific error the audit found, and it would have surfaced as a mysterious Tier B failure rate deep into Phase 2.
Decision recordwiki/adr/009-operational-identity-before-consciousness.md
251 words
Inside this chapter

ADR-009 — Study operational identity before consciousness or human continuation#

  • Status: Proposed by independent audit pass 2
  • Date: 2026-08-01

Context#

The new Wega Labs landing page frames December around consciousness, personal identity, and continuity beyond biology. The terrarium can test longitudinal synthetic-agent behavior and the causal contribution of memory, relationships, embodiment, and history. It cannot determine whether an agent has subjective experience or whether a copied human is numerically the same person.

Without a boundary, a successful persona-consistency result could be marketed as evidence for consciousness or uploading even though those conclusions do not follow.

Decision#

December adopts the claims ladder in 18: infrastructure integrity, operational agentic continuity, causal components, later human-model fidelity, and finally consciousness/substrate survival. Work may advance only one evidenced level at a time.

The initial scientific object is persistent synthetic agency. Consciousness and human continuation remain motivating questions and explicit non-claims. No human persona intake begins before a separately approved ethics and data-governance protocol.

Consequences#

  • The terrarium has publishable research goals even if the distant aspiration fails.
  • Landing-page language must not say December makes consciousness or immortality testable today.
  • Identity measures use held-out behavior, commitments, relationships, perturbation recovery, and distinctiveness—not self-report alone.
  • Human participant work becomes a separate future track with first-party consent and independent review.

Rejected alternatives#

  • Treating stable role-play as evidence of consciousness.
  • Treating behavioral similarity to a human as numerical identity with that human.
  • Removing the long-horizon motivation entirely; it is legitimate when clearly labeled as motivation rather than current evidence.
Part V

Audits and challenge record

The current verdict first, then the historical pass and reusable audit instrument.

Current audit verdictAUDIT-FINDINGS-PASS-2.md
1,606 words
Inside this chapter

December Plan — Independent Audit Findings, Pass 2#

Reviewer: Codex
Date: 2026-08-01
Scope: first audit’s changes, revised wiki, cost/parameter/determinism additions, and public lab/identity framing
Verdict: REVISE — the engineering premise remains strong; Gate 0 is not closed

Executive verdict#

December has real depth. Its strongest near-term research contribution is not a simulated civilization by itself and not a claim about immortality. It is a causal, longitudinal testbed for persistent synthetic agency: complete life histories, private information, model-independent identity state, counterfactual forks, and rigorous evaluation under change.

The previous audit made several excellent corrections—especially scale honesty, decision-level idempotency, refusal measurement, initial-condition provenance, null demographic controls, and director prompt-injection defenses. It also introduced new overclaims and overengineering. Most importantly, the new public identity framing outruns the protocol: the valley cannot establish consciousness or human survival through substrate change.

This pass adds a lab charter and claims ladder, corrects the resident-objective policy, separates three meanings of replay, demotes the parameter registry from “verified truth” to provisional evidence, corrects current pricing interpretation, and places a hard boundary around human-participant work.

Blockers#

ID Finding Consequence Resolution/status
P2-F01 The landing page said the experiment made personal continuity testable and solicited human participants, while the wiki listed conscious-being research as an anti-goal and contained no identity protocol. Brand promise and evidence were misaligned; risk of misleading users and collecting personal information without adequate governance. Resolved 2026-08-01: claims ladder and Track D boundary added in 18; both rendered implementations now lead with the living-world dream and keep consciousness/human continuation as an explicitly unproven horizon.
P2-F02 “Request a human participant” permitted nomination of another person and opened a public GitHub issue containing a name/alias and reason their life should be studied. Consent could not be verified; public personal data, family/third-party disclosure, reputational and emotional risks. Resolved for the current scope 2026-08-01: participant UI, form, and client-side issue generation removed from both implementations. December accepts fictional residents only; future human work remains blocked behind the separate protocol in 18.
P2-F03 The first audit repaired and verified its own findings in one pass, contrary to the project’s audit protocol, then declared readiness. “Independent audit complete” overstates assurance. README now records two passes and Gate 0 remains open. This report is the independent re-audit artifact.
P2-F04 The identity/continuity thesis had no registered dependent variables or baseline. Any coherent persona could be called “identity”; the flagship lab claim was unfalsifiable. Track A and the first three experiments in 18; metrics added to validation plan; a preregistration-ready R0 protocol and template added in 19. It remains a draft until frozen in version control before confirmatory runs.

High findings#

ID Finding Consequence Resolution/status
P2-F05 Replay, deterministic re-execution, and fresh counterfactual simulation were conflated. Floating-point caveats were applied to event replay unnecessarily, while live branch semantics remained unclear. 14 and ADR-006 now define three guarantees.
P2-F06 The decision barrier requires timeouts in simulated time while the simulation can be blocked waiting for the decision. Potential deadlock; provider availability silently affects whether agency exists. Live wall-clock/retry policy may choose the recorded outcome; replay consumes that recorded resolution. High-significance missing decisions pause rather than become a universal “survival” choice.
P2-F07 Per-event global lattice hashing, xid8 fencing, custom inverse-CDF distributions, and a project-owned transcendental library were placed into Phase 1 before measurements. The integrity layer risks consuming the project before a resident exists; some mechanisms solve scale December does not have. Normative minimum reduced to one serialized writer, hash-chained events, aggregate versions/hashes, canonical snapshot hashes, named RNG streams, and replay tests. Advanced mechanisms are escalation options.
P2-F08 “Generated, not authored” treats a seeded generator as neutral. Authorship is hidden in distributions, constraints, exclusions, and seed selection; emergence claims can still be planted without outcome words. Reframed as declared, randomized, and varied initial conditions; canonical exhibition separated from research cohorts.
P2-F09 The parameter registry labels rows “verified” without a source ID/locator per row and converts heterogeneous evidence into implementation prescriptions. Another reviewer cannot reproduce provenance; ecological fallacy and false precision can enter code. Registry is now explicitly provisional; source-ID/locator and transfer-justification columns are required before admission to code. Harmful prescriptive prose corrected.
P2-F10 Resident objectives still said “plural goals and survival,” and a universal survival fallback was used during outages. Suppresses martyrdom, dominance, revenge, risk-seeking, self-destruction, and other human-like variation; creates artificial peace/survival bias. Society, cognition, and operations docs now distinguish no system-wide engagement reward from heterogeneous resident motives. High-stakes cognition pauses on failure.
P2-F11 The cost table assigns GPT-5.6 Luna the old GPT-5-nano price and calls MiniMax M3’s published permanent reduction temporary. Cost comparisons and model-routing recommendations are factually wrong. Corrected against official 2026-08-01 provider pages. Model choice remains benchmark-driven.
P2-F12 A public canonical world, research cohort, and hypothesis test were treated as one artifact. Tuning the exhibit contaminates research; a single beautiful history may be presented as a result. 18 separates the canonical exhibit from registered cohorts and controls.

Medium findings#

ID Finding Resolution/status
P2-F13 The fixed founding population of 18 conflicts with the audit’s own statement that Gate 1 should determine size. Public page should call 18 the current design cohort, not a scientifically established optimum.
P2-F14 Acute epidemic “bimodality” was stated as a strict invariant using large-population approximations at n=18. Reworded as an ensemble expectation; intermediate outbreaks are possible and not automatically bugs.
P2-F15 Conflict “frequency” was proposed as an input sweep. Sweep mechanisms/dispositions/conditions; use observed conflict rates only as outputs and broad plausibility checks.
P2-F16 Historical examples and population-genetic detail exceed the needs of the first terrarium and may create false authority. Keep as research notes; genetics deferred beyond initial demography unless it becomes an explicit question.
P2-F17 The licensing table calls LGPL/GPL components “not usable,” which is too categorical and resembles legal advice. Recast as integration obligations requiring deliberate review; source reading is distinct from code reuse.
P2-F18 December was presented publicly like a mature scientific program before a research operating system, registered experiment, versioned corpus, or result existed. Public-framing portion resolved 2026-08-01: Wega Labs remains correctly identified as the company and AI lab, while December now leads as one of its long-term dreams and clearly says the world is not alive yet. 18 still defines how the project earns stronger research claims over time.
P2-F19 Two unrelated projects use the December name: the living-world/identity program and “December Sato,” an autonomous Mac mini agent with broad privileges and Twitter access. Brand confusion and a security posture incompatible with the terrarium’s containment story. Rename or clearly separate December Sato; do not describe unrestricted computer ownership as a lab method.
P2-F20 The top-level research documents are not version-controlled even though two nested website directories are separate Git repositories. Gate 0 still requires a repository containing the wiki, ADRs, audits, and experiment registrations. Owner action remains open.

What is genuinely strong#

  • The authoritative causal kernel is the correct foundation.
  • Objective events and subjective memory are cleanly separated.
  • Capability-based institutions avoid giving prose magical authority.
  • Event provenance, private information lineage, and counterfactual branches can become distinctive research infrastructure.
  • Scale honesty and null baselines are excellent additions.
  • Provider refusal, homogeneity, schema failure, and model drift are correctly treated as experimental confounds.
  • The project is unusually explicit about claims it must not make.

Documentation changes made in this pass#

  • Added 18, ADR-009, and 19: lab thesis, claims ladder, research tracks, human-data boundary, experiment-card template, and the R0 protocol.
  • Updated the README, vision, validation, roadmap, risks, audit guide/template, and landing-page brief so operational identity, canonical exhibit, and registered research cohorts are distinct concepts.
  • Corrected cognition and conflict policy so the system has no observer-engagement reward while individual residents may still value dominance, territory, revenge, glory, risk, altruism, or self-sacrifice.
  • Split recorded-event replay, pinned kernel re-execution, and fresh counterfactual branching; replaced Phase 1’s advanced hashing/fencing mandate with a serialized writer and proportional integrity checks.
  • Recast initial conditions as authored but declared/randomized/varied, with proxy-feature and seed-selection audits.
  • Demoted the parameter registry to a provisional research notebook and corrected prescriptive or scale-inappropriate claims, including strict epidemic bimodality and conflict-rate tuning.
  • Corrected current model-price interpretation and clarified that schema transport validity is not semantic/action validity.
  • Recast reciprocal-license components as review decisions rather than categorically unusable software.
  • Initially marked the participant intake as a release blocker, then removed it from both rendered implementations in the follow-up landing-page copy pass. Calls to action now invite questions and project ideas, not human-life data.
  • Marked the first audit historical so its superseded remedies cannot be mistaken for the current specification.

Gate 0 status#

Gate 0 remains open until:

  1. the five owner decisions in the first audit are supplied;
  2. the research corpus is under version control;
  3. ADR-009 and the claims ladder are accepted or revised;
  4. public participant solicitation is disabled or replaced — resolved 2026-08-01;
  5. a first registered continuity experiment and experiment-card template exist;
  6. parameter rows used by Phase 1 have reproducible source locators and transfer justifications;
  7. the Phase 1 integrity minimum is accepted, with advanced machinery deferred until measurement requires it.

Final assessment#

Proceed—but proceed as a lab that happens to have a beautiful long-horizon vision, not as a vision looking for scientific language. If the first year produces a credible continuity benchmark, a replayable causal-agent testbed, and two careful negative or positive studies, Wega Labs will have earned the label. If it produces only a striking website and one cinematic valley history, it will not.

Historical auditAUDIT-FINDINGS.md
2,846 words
Inside this chapter

December Plan — Audit Findings, Pass 1 (historical)#

Status notice (2026-08-01): This report records the first audit and is not the current gate verdict. Several remedies below were themselves revised after a second independent audit—for example strict epidemic bimodality, xid8 fencing, per-event lattice hashing, “generated not authored,” and cost/pricing claims. Use AUDIT-FINDINGS-PASS-2.md for the current verdict and the wiki/ADRs as the current specification.

Reviewer: Claude (Fable 5), adversarial audit pass Date: 2026-08-01 Plan version: pre-git (no repository initialized — see F-19) Verdict: REVISE → resolved in place. READY FOR PHASE 1 conditional on the five owner decisions in §Open.

This pass followed wiki/12-audit-guide.md. Unlike the protocol's default, findings were remediated in the same pass rather than left for a response document, because the plan had not yet been implemented and the corrections were structural rather than contested. Every finding below states where the fix landed.

Executive summary#

The plan is unusually good. Its instincts — kernel owns truth, event-sourced history, capability-based authority, read-only director, no arbitrary agent code, ablations over anecdotes — are the right ones, and it anticipates most of the ways this kind of project fails. The writing is disciplined about the difference between what is designed and what is claimed.

Its weakness is uniform and diagnosable: it deferred every question whose answer might constrain the ambition. Parameters were "to be calibrated in Phase 1." Cost was "to be refreshed before budgeting." Determinism was asserted rather than specified. Initial conditions were never assigned an owner. Four of those deferrals turned out to hide problems that would have surfaced deep into implementation, when they would have been expensive or fatal:

  1. The replay guarantee, on which every integrity promise depends, was not achievable as written.
  2. The disease module was specified around a mechanism that cannot operate at eighteen people.
  3. The founding scenario contained the plot it was supposed to discover.
  4. The cost model — the top-rated operational risk — had no bottom-up estimate, and the one figure it did cite priced a promotional rate for a now-legacy model.

A fifth issue is subtler and, in my judgment, the most likely to embarrass the project later: at n=18, demographic noise will produce collapses that look exactly like causal cascades, and nothing in the validation plan distinguished them.

All five are now addressed. Four new documents were added (14, 15, 16, 17), three new ADRs recorded, fifteen risks added, and the bibliography corrected — it contained a fork cited as an upstream project, two wrong licenses, a misattributed failure taxonomy, and several stale facts.

The premise survives the audit. Nothing found here makes the core promise impossible. Two findings (F-05, F-06) do constrain what the world can honestly claim to be, and one (F-04) means the settlement needs its neighbors earlier than planned.

Blocker findings#

ID Finding Consequence Resolution
F-01 "Bit-for-bit identical replay" is unachievable for a Python float kernel. glibc explicitly disclaims correctly rounded transcendentals; FMA contraction and SIMD reduction order differ across architectures; NumPy guarantees nothing across builds. Separately, no LLM provider offers reproducible sampling — Anthropic has no seed parameter at all. Every integrity promise — event sourcing, shadow worlds, causal UI, the audit protocol itself — rests on a claim that would fail the first time an auditor replayed a history on their own machine. 14 + ADR-006. Restated as pinned-environment determinism, with an integer-state option decided at Gate 1. Response cache made canonical; replay hard-fails on a miss.
F-02 Reading the event log by bigserial high-water mark silently loses events. Sequence values are allocated at insert but transactions commit out of order, so a projection can advance past an event that has not yet committed and never see it. Rollbacks leave indistinguishable gaps. Permanent, silent projection corruption. The only check that would catch it is the rebuild-equals-live invariant, and by then the cause is long gone. 14 §D7. Enforced single writer under advisory lock, plus xid8 snapshot fencing as defence in depth. Sequence-gap injection test added.
F-03 The founding scenario pre-installs its own outcomes: founders who "disagree about property, leadership, risk"; land claims that are "ambiguous"; a charter whose assembly procedure is "unsettled." No document assigned ownership of initial-condition generation. The project's central claim collapses. A settlement that fractures over property and elects a leader would be reproducing what was planted, and the causal graph would be honest and worthless — the first cause sits outside recorded history. 17 + ADR-007. Scenario rewritten to material facts only; seeded generator; outcome-label scan in CI; mandatory no-charter control arm; pre-registered definitions.
F-04 Eighteen people is below every modelled viability threshold (demographic ABMs put the floor at 40 and near-certainty at 150), the founder lineages exhaust within two to three generations, and demographic variance rivals the mean. Outsiders were deferred to Phase 5. Every Phase 3 and Phase 4 institutional experiment would run inside a closed, guaranteed-declining population, confounding all institutional findings. 15 §C, 01, 03. Aggregate neighbor moved to Phase 3; exogamy and in-migration promoted to load-bearing mechanics; a regional aggregation added because one same-sized neighbor is insufficient; founding size to be set by Gate 1's extinction curve.

High findings#

ID Finding Consequence Resolution
F-05 The disease module cannot work as specified. Critical community size for acute directly-transmitted immunizing infections is 250,000–500,000. December has 18 people. A generic SEIR pathogen will either contribute nothing or, if tuned until epidemics appear, encode a rate that cannot physically exist. Fabricated epidemiology presented as a validated subsystem — risk R-07 in scientific costume. 03, 15 §D. Replaced with a five-mechanism model (introduced epidemics, environmental, zoonotic, chronic/latent, helminth). Epidemics required to be bimodal. Realism contract gains R7, scale honesty.
F-06 Small-population noise is indistinguishable from causal emergence. With expected births and deaths in single digits per decade, many runs end for no interesting reason. No baseline existed. Collapse narratives would be reported as cascades when they were arithmetic. This is the finding most likely to survive into a published claim and then be demolished. 09. Null demographic model made a prerequisite; all cascades reported as differences from it. Risk R-29.
F-07 No bottom-up cost model for the top-rated operational risk. The only figure cited rescaled another project's token volume at a promotional MiniMax rate for a model now legacy. The project could not answer whether it is affordable, and would have discovered the answer while running. 16. Built from activation counts: $50–$1,300/month, cross-checked against Agentopia to within a factor of ~2.
F-08 Tier C dominates cost — ~10% of calls, ~89% of spend. The plan treated significance routing as a quality mechanism, not the primary budget lever. Optimization effort aimed at the wrong target; population size wrongly perceived as the cost driver. 16 §4, 04. Significance threshold given live tuning, dashboard, and alert.
F-09 Sparse activation defeats prompt caching, and the conflict was unnoticed. Input is >90% of tokens; cache misses cost ~10×; TTLs are 5–30 minutes while activation is deliberately scattered. An architecture optimized for fewer calls can cost more than one optimized for cache locality — the cost model inverts. 16 §4, 07. Cache-aware activation batching; TTL as a model-selection criterion; hit rate as a primary metric; Gate 2 threshold ≥60%.
F-10 The affordable models cannot honor the schema contract. MiniMax's OpenAI-compatible endpoint does not support response_format at all; DeepSeek and Qwen are best-effort only. The plan assumed MiniMax could carry routine cognition. Schema failures concentrated exactly on the tier carrying the volume, discovered deep in Phase 2. 16 §5 + ADR-008. Validity ladder with escalation to a schema-guaranteed model — the same pattern Agentopia used.
F-11 Model refusal was an unhandled failure mode. Providers may decline to represent raiding, deception, or household formation. Not merely operational: if a safety layer makes agents reluctant to escalate, December reports "peaceful institutions emerged from material conditions" when the cause was the provider's training. 16 §6, 04, 08. Refusal benchmark in admission; distinct failure class; rates reported with every behavioral claim.
F-12 Per-event full-state hashing is O(state) per event, making the pipeline quadratic in world size. At 10⁴–10⁶ entities this becomes the dominant cost of the entire simulation. 14 §D6. Incremental lattice hashing (O(changed entities), order-independent), periodic full verification, Merkle at snapshots for divergence localization.
F-13 Decision-level idempotency was missing. Command idempotency keys stop duplicate delivery, not a timed-out call whose retry returns a different command. Two valid non-duplicate commands for one decision; history forks silently from the replay. 14 §D4, 06. decision_id is the idempotency unit, enforced by unique constraint.
F-14 Rejection responses leak hidden state. Returning "feasible alternatives and reasons" can reveal a granary's contents or an unobserved occupancy — a direct violation of realism contract R3. Partial observability quietly broken through a channel the canary suite did not probe. 06, 08. Visibility filter on all rejection responses; canary probes of the rejection channel added.
F-15 Behavioral homogeneity. Twelve residents on one model with one template converge; the plan varied inputs without guaranteeing varied outputs. No factions, uniform voice, and any factions that do appear are artifacts of the value draw rather than social process. 04. Diversity metrics, per-resident model/prompt variation, homogeneity check in admission. Risk R-31.
F-16 The director is an unguarded prompt-injection target — resident text flows into the summarizer whose output reaches the human. Also unescaped in the UI. World state stays intact while the observer's understanding is corrupted, with the read-only guarantee technically satisfied throughout. 07, 08, 09. Delimited untrusted data, citation-enforced factuality checking, UI escaping, dedicated red-team scenario.

Medium findings#

ID Finding Resolution
F-17 Bibliography errors. GovSimElect is a five-star personal fork cited as an upstream project (real work: giorgiopiatti/GovSim, NeurIPS 2024). Craftium/Luanti are LGPL-2.1+ with CC BY-SA media, not permissive. Agentopia has no LICENSE file despite a README claim of MIT. project-sid contains no code. MemFail's actual taxonomy is summary/storage/retrieval/reasoning failure, not the list quoted. Melting Pot's substrate counts come from the README, not the abstract. Mindcraft moved owners; Covasim moved to starsimhub; Unknown Horizons is dormant since 2019; the OpenRouter docs URL 404s. 13 rewritten; every URL re-verified; 02 candidate matrix corrected.
F-18 Mesa facts were stale. Current stable is 3.5.1; event scheduling is now stable, not experimental; mesa.experimental.devs is deprecated and removed in 4.0; mesa.space is maintenance-only. 02, 13. Prior revised toward Mesa as a toolkit rather than the kernel.
F-19 No version control. The audit template asks for a plan commit, the event envelope records code_version: git:..., and pre-registration depends on provable history. Added to Gate 0. Do this first.
F-20 Event volume was never estimated while committing to indefinite retention. Grid resolution is the dominant driver — one event per cell per day at 50 m resolution is millions of events and gigabytes per simulated year from vegetation alone. 14 §D8, 06. Batched world events; ≤5,000 events/sim-day target; time partitioning; cold-tier archival.
F-21 Gates were unfalsifiable — "large safety margin," "cost targets pass." 10. Gates 1, 2, 3, and 5 quantified.
F-22 Phase 1 estimate was optimistic at 3–6 weeks for event sourcing, replay, ecology, hydrology, crops, demography, and an experiment harness — before the determinism work this audit added. 10. Revised to 6–12 weeks with the reasoning stated.
F-23 Success criteria were unfalsifiable. "Three distinct governance forms" had no definition of distinctness. 00, 17. Pre-registered structural definitions required.
F-24 Violence rates treated as calibratable. The empirical literature genuinely disagrees, by an order of magnitude, about warfare mortality in small-scale societies. 15 §E. Represented as a swept parameter across a contested range; strengthened the claims-we-will-not-make list.
F-25 Licensing question was open but answerable. 13. Decision table added. MiroFish's AGPL-3.0 is the one genuine trap, since its obligations trigger on network use.
F-26 Python version guidance was behind (3.13+ stated; 3.14 is current, 3.15 due Oct 2026), with no position on free-threading. 06. Default GIL build required for the kernel.

Second research pass — a partial self-correction#

The first pass exhausted its budget on epidemiology and left sections A, B, C, E, and F of the parameter registry unverified. A second pass closed them, and it revised one of this audit's own findings.

F-04 was overstated. The first draft asserted that eighteen people "cannot persist multigenerationally." The models do say that — but the historical record contains counterexamples the audit had not checked: Pitcairn persisted from 27 founders, Tristan da Cunha from about 15, and Rapa Nui recovered from a nadir of ~110. The corrected claim is that an 18-person founding is marginal rather than impossible, and the design consequence is unchanged but better grounded: what separates the survivors from Norse Greenland, Roanoke, and the Polynesian Henderson colony is network connection, not headcount.

That correction improved the plan rather than weakening the finding. It also surfaced four historical mechanisms now specified in 15 §C-4 — founding sex ratio driving lineage-erasing violence (Pitcairn lost every Polynesian male line within a decade), correlated-risk cohort loss (Tristan lost 79% of its adult men in one boat in one day), demographic damage outliving its cause (St Kilda's tetanus ended in the 1920s; the island evacuated in 1930 regardless), and exchange networks carrying marriage partners alongside goods.

Other corrections the second pass forced on the audit's own provisional numbers: walking energy cost is 0.81 kcal/kg/km gross, not 0.5; storage loss is ~3% for storage proper against a 13% whole-chain figure, not 18%; and forager and horticulturalist fertility are statistically indistinguishable (5.6 vs 5.4, p=0.8) rather than differing as the first draft implied. The seed-retention estimate held up almost exactly — 27% verified against 28% estimated, now computed from 26,000 medieval manor-year observations rather than recalled.

What the audit did not find#

Stated explicitly, because a finding list reads as a verdict on quality and this one would mislead:

  • No hidden narrative mutation path. The director is genuinely read-only, and the separation of prose from state is maintained consistently throughout.
  • No security architecture gaps beyond F-14 and F-16. ADR-005 is correct and well-argued.
  • No confusion about authoritative truth. ADR-001 is the single best decision in the plan and it is applied consistently.
  • No ethical blind spots in the interpretive limits — the anti-goals and the "claims we will not make" list are stronger than most published work in this area.
  • The realism-tier system works. It is the mechanism that keeps ambition from becoming scope creep, and it should be defended in later phases.

Open — requires the owner, not the auditor#

These block Gate 0 and cannot be resolved by analysis:

  1. Monthly spend cap and emergency cutoff behavior. 16 supplies the numbers to decide against; the decision is yours.
  2. Licensing posture — permissive-only or GPL-compatible. 13 now states what each choice costs.
  3. Deployment target — machine, OS, and whether cloud hosting is in scope. This interacts with F-01: a single pinned platform makes the determinism claim much easier to keep.
  4. Canonical pace range, which drives cost roughly linearly.
  5. Publication intent — private, or a public history. This changes the AGPL analysis, the privacy posture, and how carefully the "claims we will not make" list must be policed.

Re-audit checklist#

  • [x] Every Blocker finding resolved and verified.
  • [x] Every High finding resolved and verified.
  • [x] Medium findings dispositioned with a target gate.
  • [x] Changes introduce no new contradiction in the traceability matrix.
  • [x] Known-limit language remains honest — and is now stronger, since the determinism, epidemiology, and violence sections all state what cannot be claimed.
  • [ ] Owner has supplied the five decisions above.
  • [ ] Version control initialized (F-19).
  1. git init, commit the plan as it stands, and tag it. Everything about provenance depends on this and it takes a minute.
  2. Answer the five owner decisions.
  3. Re-verify the provenance-class L parameters in 15 §§A, B, C, E, F — this audit exhausted its research budget on the disease section, which is the one that changed the design.
  4. Begin Phase 1 with the integer-versus-float state decision (ADR-006) as the first task, because it cannot be retrofitted.
  5. Build the null demographic model before any social mechanism, so every later claim has a baseline.
Review instrumentAUDIT-REPORT-TEMPLATE.md
486 words
Inside this chapter

December Plan — Independent Audit Report#

Reviewer:
Date:
Plan commit/version:
Verdict: BLOCK / REVISE / READY FOR PHASE 1

Copy this file to AUDIT-REPORT.md. Do not edit the plan during the first audit pass. Follow wiki/12-audit-guide.md.

Executive summary#

Summarize whether the proposed system can produce persistent, causally grounded, observable emergence within its stated scope and operational constraints.

Findings#

ID Severity Area Evidence (file/section) Consequence Required remediation Status
A-001 Blocker/High/Medium/Low Open

Core-promise verdicts#

Promise Supported? Evidence Gap/finding ID
Materially grounded survival Yes/Partial/No
Residents build novel things Yes/Partial/No
Institutions and elections Yes/Partial/No
Factions, diplomacy, and war Yes/Partial/No
Disease, disaster, and extinction Yes/Partial/No
Unscripted but explainable emergence Yes/Partial/No
Return-after-two-days observer experience Yes/Partial/No
Continuous operation within bounded cost Yes/Partial/No
Operational identity is measurable beyond style/self-report Yes/Partial/No
Public claims remain within the evidence ladder Yes/Partial/No
Canonical exhibit is separated from confirmatory cohorts Yes/Partial/No

Causal-integrity review#

  • Hidden narrative mutation paths:
  • Missing invariants:
  • Unconditioned or unsourced randomness:
  • Causal graph/replay gaps:
  • Partial-observability leaks:

Realism and model-validity review#

  • Unsupported scientific/historical claims:
  • Parameter-provenance gaps:
  • Structural-realism concerns:
  • Missing patterns/ablations:
  • Claims that exceed fitness for purpose:

Agent and memory review#

  • Autonomy versus kernel control:
  • Identity continuity:
  • Memory/source lineage:
  • Private information:
  • Tiering/model-swap risks:

Research-program and claims review#

  • Claim-ladder classification of every headline claim:
  • Registered dependent variables, baselines, falsifiers, and stopping rules:
  • Style, rigidity, prompt-length, provider, and task-leakage confounds:
  • Canonical exhibit versus research-cohort separation:
  • Negative-result and replication path:
  • Evidence that “lab” describes operating practice rather than branding:

Economy, construction, and institutions review#

  • Stock/flow and claims:
  • Novel-project boundary:
  • Institutional grammar/capabilities:
  • Election lifecycle:
  • Failure, refusal, enforcement, and legitimacy:

Conflict, hazards, and ethics review#

  • Escalation/logistics/aftermath:
  • Peaceful alternatives and non-glamorization:
  • Disease/fire validity:
  • Extinction classification:
  • Content/interpretive risks:
  • Human/third-party data intake surfaces:
  • Consent, review, withdrawal/deletion, retention, and incident controls:
  • Consciousness/personhood/human-continuation overclaims:

Architecture, operations, cost, and security review#

  • Event sourcing and transactions:
  • Determinism and replay:
  • Provider/budget failure:
  • Backup/recovery/soak readiness:
  • Prompt injection, arbitrary tools, secrets, and authorization:

Traceability gaps#

List promises lacking any of: authoritative state, mechanism, evidence/UI, validation test, implementation gate.

Promise Missing link Required addition

Unresolved assumptions and questions#

Required changes before re-audit#

Re-audit checklist#

  • [ ] Every Blocker finding is resolved and verified.
  • [ ] Every High finding is resolved and verified.
  • [ ] Medium/Low findings have an owner-visible disposition and target gate.
  • [ ] Changes introduce no new contradiction in the traceability matrix.
  • [ ] The known-limit language remains honest.
  • [ ] Public pages contain no unreviewed human-participant solicitation or third-party nomination.
  • [ ] Operational identity is not presented as consciousness, personhood, or human survival.
  • [ ] Confirmatory claims come from frozen cohorts rather than the canonical exhibit.
  • [ ] Owner has supplied spend, licensing, deployment, pace, and privacy decisions needed before Phase 1.

Final sign-off#

Reviewer verdict after remediation:
Open Blockers:
Open Highs:
Conditions attached to readiness:

Reproducibility

Corpus manifest

Re-run ruby scripts/build_book.rb after changing a source document.

Generated: 2026-08-02T07:04:44+05:30 · Source revision: fbee07e · Corpus SHA-256: 7c02fc69737a228c02c58008cbec27e2fce32d81a675bacdc31f1099a53a6d0d

SourceWordsSHA-256
OVERVIEW.md2837706eb714f9d7
README.md1141802646fe93ec
LANDING-PAGE-BRIEF.md7395db1647a4d17
wiki/00-vision-and-north-star.md1113c2a5554bbe28
wiki/01-scope-and-realism-contract.md1649a4e0125e9179
wiki/02-research-landscape.md152321fe09863ca3
wiki/03-world-model-and-scenario.md23386ca0f65e4ed8
wiki/04-agents-cognition-and-memory.md1761723ef1fa1ea8
wiki/05-society-economy-governance-conflict.md151553a25e80d808
wiki/06-architecture-and-data.md1642129756947ac7
wiki/07-time-emergence-and-observation.md1427a369c5e176b8
wiki/08-models-cost-operations-security.md1617813ff0906b37
wiki/09-validation-and-experiments.md1809fed40ef7e4b9
wiki/10-roadmap-and-gates.md18316f1af92b0534
wiki/11-risks-decisions-open-questions.md2433306ff8203cca
wiki/12-audit-guide.md14384f6942d279f0
wiki/13-sources.md3526c9d5650bc2ca
wiki/14-determinism-replay-and-state-integrity.md332068cf71431be4
wiki/15-parameter-registry.md84355ac62aed6f72
wiki/16-cost-model-and-model-selection.md34816b676bd2cabe
wiki/17-initial-conditions-and-authorship.md1473fe1ed39c2e2c
wiki/18-lab-charter-and-research-program.md17971e94a4a2d4e8
wiki/19-experiment-card-template-and-r0-protocol.md1565eb722c83391e
wiki/adr/001-causal-kernel-not-llm-world.md14495c2f4329b4e
wiki/adr/002-event-sourced-history.md260938d33bcb26a
wiki/adr/003-replaceable-world-adapter.md141c07557458fe1
wiki/adr/004-fictional-early-agrarian-start.md4421d9a8e6440b6
wiki/adr/005-no-arbitrary-agent-code.md139444fe1649b9d
wiki/adr/006-pinned-environment-determinism.md6871b2d7d166545
wiki/adr/007-generated-not-authored-initial-conditions.md44937df5e44af33
wiki/adr/008-cheap-model-first-with-escalation.md629de96bf46c312
wiki/adr/009-operational-identity-before-consciousness.md25186ad0e58ae16
AUDIT-FINDINGS-PASS-2.md1606235956316aef
AUDIT-FINDINGS.md2846e6fe9ea86309
AUDIT-REPORT-TEMPLATE.md486db56d75985a4