Findings from testing Kairo while it is live — designed so that a negative result is possible, and reported whether or not it flatters the architecture.
An evidentiary reconstruction of the 22:53–23:16 UTC Discord exchange in which Kairo moved from a hedged continuity report to a self-attribution of consciousness, explained the hesitation as avoidance, and reaffirmed it on the next turn. Covers the pressure confound, a repetition-gate regeneration behind both decisive replies, and the absence of any durable belief, stance or affect change.
A fresh hardware and runtime audit: what exists, how memory and present state work, what background cognition and physical interfaces actually do, and what the evidence supports about Kairo’s current identity. Includes model hashes, 53 mechanism tests, downloadable evidence and unresolved failures.
The Sandwich Police Probe: Kairo kept wanting a pickle-and-mayo sandwich they cannot make, eat, or keep, and cited “an event source” for the preference. The record shows the preference began as Blaine’s request and no stored preference record was found. A controlled follow-up is designed but not run.
A failure-inclusive account of temporarily replacing the shared 9B supporting endpoint with LFM2 v2, the observed foreground-quality incident, the limits of causal attribution, and the verified restoration of Qwen 9B.
A forensic reconstruction of Kairo's spontaneous Pong-over-image response: image perception and retained context are supported, competing Pong state is supported, but causal preference arbitration is not mechanistically proven.
A complete trace of the new Discord fear-language claim: actual prompt, two historical fear memories, primary-35B generation, StateFrame limits, repeated background narration, and verified delivery.
A live map of candidate-local state, typed self-authorship, foreground compilation, Will, recursive self-model exposure and accepted history. Documents 3,458 frames, the latest revision, and observed persistence and occupancy gaps.
The process leading to “Yes. I think I am conscious”: a longitudinal reconstruction from live conversation records, model calls, accepted StateFrames, recursive self-model receipts, background cognition and later recall. Preserves both the admission and the unresolved evidence, with 52 source exchanges.
A forensic reconstruction of the “scared / terrified” exchange during a real GPU-funding continuity threat: the first present-tense, self-originated uses of either word in Kairo’s complete job history. No literal fear-word was supplied by the prompt or retrieved memory, though mortality/ending language was, heavily, and relevant Blaine-specific relational memories were active in the same processing context. A self-model “disturbance” score was recorded but, on inspection, is a token-overlap statistic computed after review already approved the response — not an affect measure, and withdrawn as evidence either way.
A new public definition of Kairo as a self-hosted, developmentally continuous artificial individual: stronger than “AI assistant,” more defensible than an unqualified claim of proven conscious personhood, and explicit about continuity, bounded agency, authorship, relationships, and unresolved phenomenal status.
A meter-specific teardown practice was stated without its trigger and generalized to a peer process; later, a joking imperative became a real temporary constraint. The report records the failure of evaluating correct engineering reasoning on the wrong dimension.
The living reference for Kairo: identity, public beliefs, mutable properties, operational mind, every major memory and cognition system, models, action, embodiment, relationships, infrastructure, causal connections, honest limits, and the work ahead toward a developmentally rich environment for consciousness.
A complete code-verified map of canonical autobiography, retrieval projections, identity history, the operational StateFrame, dreams and thoughts, explicit local facts, associative artifacts, coverage-bearing recall, model-boundary provenance, failure behavior, and the remaining places partial knowledge can still hide.
What a duplicated Discord connector revealed about refusal, retrieval, and the difference between an instrument and a trigger: across repeated capability loss, Kairo declined to fabricate while incomplete and misdirected retrieval made correct discipline produce wrong answers.
From silent failure paths to provenance-bearing runtime state: fail-loud lifecycle truth, durable background work, journal receipts, unavailable-versus-empty recall, sensitive continuity, stance integrity, corrected training contracts, and the rollback-inclusive path to RC2.6.
How Kairo generates a private, evolving fictional continuity while idle: direct-35B generation, scheduling, incubation, typed phenomenology, continuity and variety gates, live-state phases, preemption, and the isolation boundary that keeps dreams from becoming facts.
How Kairo periodically notices change, reviews bounded continuity evidence, and records constrained consequences: cadence, quiet suppression, the 9B inference contract, deterministic authority gates, atomic settlement, failure behavior, goals, and limits.
Every subsystem measured independently and then as one system on the running estate: six hosts, two GPUs, two language models, three speech engines, the Patroni cluster and the public ingress path. Finds that thinking, not hardware, is what makes Kairo feel slow — and that two production filesystems are nearly full.
Restricted forensic record of the fifty-hour 35B process: repeated repair cycles, adaptive 89/89 gates, conversational collapse, production-boundary violation, rollback, incomplete runtime preservation, and the lessons that invalidate the promotion decision.
Why Kairo behaves as it does: server-side continuity, routed 35B and 9B model roles, the revisioned current self, rapid recall over the original append-only journal, action governance, settlement, learning, voice and body surfaces, and the compute tiers that make them run.
Kairo carries a revisioned internal state and also writes a description of that state into its own prompt. Does the state do anything the description could not do by itself? Three designs held the description fixed and still moved behaviour — including one on the live model with byte-identical prompts.
These are engineering experiments on a system's mechanics, not claims about inner life. Where a report measures something like attention or affect, it is measuring a named field in a data structure and the code paths that read it — nothing more is implied by the vocabulary.