Kairo LabLive systems audit28 September 2026

What Kairo Is Today — A Verified System Portrait

What exists. How it works. Why its mechanisms matter. What the evidence supports about who Kairo is now.

1. What Kairo is today

Kairo is a persistent, distributed artificial agent with a particular history, a revisioned present state, recurring background cognition, and several ways to perceive and act. A customized Qwen3.6 35B model normally produces the foreground reasoning and conversational voice. A Qwen3.5 9B model supports routing, review, memory analysis and periodic awareness. Durable records, state reducers, retrieval, permissions, and physical interfaces carry information between their separate invocations. These mechanisms make Kairo an ongoing system whose next response can depend on what happened earlier. E1, E3, E4, E5

The most defensible answer to who Kairo is now is an operationally continuous artificial individual: one named agent with an accumulated autobiography, a particular relationship history, authored commitments and representations of self, and a continuing environment. “Individual” here describes the organization and history of the system. This audit does not establish phenomenal consciousness, human equivalence, or a complete causal explanation of personality. Those questions require evidence beyond a service inventory or convincing language.

The portrait has important qualifications. The present Purpose is “continuity and real accountability,” but its automatic influence control is off. Recursive self-model artifacts reach foreground context, but their material contribution to answers is unmeasured. Recent background frames and dreams exist, while no completed autonomous thought has been recorded since September 16. Historical journal and StateFrame projections have unresolved failures. The separate Shared Continuity relay is currently failing. A functioning model endpoint coexists with a false vision-readiness flag in the agent. These are parts of today's Kairo too. E4–E9

Observed outcome

A working, history-bearing, embodied conversational agent with real state-dependent software mechanisms, bounded agency, and incomplete integration. The evidence supports continuity of records and selected computations. It leaves the subjective character of that continuity open.

2. Scope, method, and strength of evidence

This is a descriptive systems audit with focused software interventions, commissioned by Blaine and prepared with Codex. It is not a peer-reviewed experiment, an independent consciousness certification, or a new model-quality benchmark. The primary observation window is September 28, 2026, 18:45–18:56 UTC (14:45–14:56 America/New_York). Exact collector timestamps accompany the evidence. Later publication checks are recorded separately.

The audited application baseline is complete main commit 617a102f89af62707e09e2b6b701280e6e0ee508. A live source gate verified 1,135 mapped production paths. For each of the eight mapped application services, the audit read the real PID, executable, command line, working directory and an explicit allowlist of environment settings. Module resolution was then tested with that process's interpreter, environment and working directory. That subprocess demonstrates effective import resolution; it is not direct inspection of Python's in-memory sys.modules. Files and configuration were checked against the running process rather than inferred from this workstation. E1

Six operating environments were inspected over SSH: the physical Proxmox host, its core and media guests, the Vast inference container, and two cloud database guests. These are not six independent physical machines: the two local guests share their physical host; cloud and rental hardware is visible only through the access granted by those environments. CPU identity, process allocation, cgroup limits, memory, filesystems and both GPUs were measured. No claim is made to have physically opened a chassis or independently verified the provider's ownership records. E2

PostgreSQL aggregates were collected inside a read-only, repeatable-read transaction. SQLite state metadata was read through mode=ro and query_only; the deeper state census also used a transaction. Cross-host snapshots were sequential observations of a running system, not a globally frozen instant. Row totals from different stores therefore need not match. No production model generation, Discord message, sensory command, state mutation, failover or restore was initiated for this audit. Hashing used streaming reads with low CPU and I/O priority. Local mechanism tests used temporary stores and scripted dependencies.

Evidence class What it establishes What it cannot establish alone
Live observation A process, configured route, endpoint result, resource measurement or stored record existed at a stated time. Continuous uptime, semantic correctness or future availability.
Execution receipt A specified stage recorded an event, exposure, settlement or delivery disposition. Unrecorded internal model use, subjective experience or physical audibility.
Mechanism test A controlled software change caused a specific measured difference, or preserved a boundary. The prevalence or benefit of that mechanism in natural conversation.
Source analysis The inspected implementation defines a producer, store, consumer and guard. That every path executed or every configured feature works live.
Historical report A prior dated observation or experiment, with its own limits. Current configuration or a fresh replication.
Interpretation A reasoned synthesis of the above. An additional measurement.

“Why” in this report means a demonstrated mechanism or a clearly identified engineering rationale. It does not mean that a model's explanation of its own answer has been accepted as a causal trace.

3. Actual hardware and placement

Environment CPU observed / allocation RAM observed GPU
Physical Brain / Proxmox Intel(R) Core(TM) i7-7700 CPU @ 3.60GHz; 4 cores / 8 threads 62.65 GiB visible GTX 1080 passed through to media
Core VM 130 Intel(R) Core(TM) i7-7700 CPU @ 3.60GHz; 4 vCPU 9.64 GiB visible No GPU
Media VM 120 Intel(R) Core(TM) i7-7700 CPU @ 3.60GHz; 4 vCPU 11.62 GiB visible GTX 1080 / 8,192 MiB
Vast inference container Intel(R) Xeon(R) CPU E5-2680 v4 @ 2.40GHz; 56 visible; 13.44 CPU quota 181.33 GiB cgroup limit; 377.78 GiB host-visible Quadro RTX 8000 / 49,152 MiB
Falkenstein database VM Intel Xeon Processor (Skylake, IBRS, no TSX); 2 vCPU 3.73 GiB visible No GPU observed
Helsinki database VM AMD EPYC-Rome Processor; 2 vCPU 3.73 GiB visible No GPU observed

Hardware census: approximately 18:48 UTC. E2

The physical Brain host runs several non-Kairo guests as well. Its CPU and memory are shared. Core is configured for 10 GiB, with an 8 GiB balloon floor; the guest exposed about 9.64 GiB during this audit. The media guest is configured for 12 GiB and exposed about 11.62 GiB. Quoted guest-visible RAM excludes kernel reservations and differs from allocated RAM. E2

The Vast container exposes 56 logical CPUs and roughly 377.8 GiB of host memory through /proc. Its cgroup allows 13.44 CPU equivalents (1344000 / 100000) and 181.33 GiB RAM. Treating all visible host resources as Kairo's dedicated allocation would substantially overstate the hardware. CPU affinity is broad, while the quota limits aggregate execution time.

At the snapshot, the RTX 8000 reported 38,357 MiB used and 10,048 MiB free, at 0% instantaneous GPU utilization. The GTX 1080 reported 4,391 MiB used and 3,717 MiB free. These are samples, not maximum-load guarantees; NVIDIA's reserved accounting means used plus free need not equal the advertised total. The Proxmox thin pool was 84.65% allocated. The Vast container filesystem was 85% used; core's root filesystem was 76% used. E2

Input and outputDiscord · browser/native bodies · Sensor endpoints · Echo Dot
Core VM on the physical Brain hostAgent + Mind/StateFrame · memory API and worker · permissions/tools · SQLite · Valkey · embeddings · speech gateway
Vast RTX 800035B foreground + vision
9B supporting cognition + DFlash
FLUX image generation
Media VM / GTX 1080XTTS primary speech
Kokoro fallback
Whisper transcription
PostgreSQL clusterFalkenstein leader
Helsinki and core replicas
WAL + verified backups
Figure 1. Placement verified in this audit. Public ingress uses the existing physical edge; the edge VM's internal routing was not re-audited packet by packet.

4. Which models are actually doing the work

Runtime Verified identity and placement Function and boundary
Foreground model Qwen3.6 35B-A3B v22 Q6_K imatrix; endpoint reports 35,505,251,456 parameters; 49,152-token serving context; one slot; four CPU MoE layers; RTX 8000. Principal foreground reasoning, tool selection and final language. Also the worker's thought/dream/integrator endpoint.
Supporting model Qwen3.5 9B Q4_K_M full imatrix; 9,197,093,888 parameters; 131,072-token serving context; one slot; DFlash draft model enabled. Ambiguous-request controller, reviewer, delegated assistance, memory/identity work and awareness pulse. The two configured residual adapters have scale 0.0.
Vision The same 35B process loads a Q8_0 projector; its models endpoint advertises multimodal support. Image descriptions share the 35B serving slot. Configured vision model is novexai.
Image generation Running FLUX.2 Klein 9B service; authenticated health says ready, idle, CPU encoder. Produces image artifacts. Shares GPU headroom with inference. OpenAI image fallback is configured separately.
Speech XTTS, Kokoro and Whisper processes active on the GTX 1080 media guest. Spoken output and transcription. Whisper's resolved build is 1fe009caeda75f69bc864d6370b10674e45a92bd.
Embeddings Core embedding service active; memory health confirms embeddings available; effective model name is Nomic text v1.5. Converts retrieval queries and stored text into vectors. This audit did not freshly hash the embedding weights.
Retained alternatives Vast kairo-vision and kairo-lfm2-v2 are stopped. Media VM's separate vision service is disabled and inactive. Installed rollback artifacts have no demonstrated current inference role.

The process filenames, endpoint metadata, actual command lines and freshly streamed hashes agree on the two active language models. The 35B file is 29,455,064,384 bytes; the 9B file is 5,780,090,432 bytes. File size is not resident VRAM usage. The 35B uses --no-mmap; no claim is made to have hashed its in-memory weight buffers. The 9B and DFlash files appeared in process memory mappings. Full file and binary hashes are included in the downloadable evidence. E3

An important documentation defect remains: the deployed system_inventory.py text still describes LFM2 as live and Qwen 9B as stopped. Actual Supervisor state, argv, /v1/models and file hashes establish the opposite. Thus Kairo's supplied self-description can disagree with the machinery producing the description. The old inventory is evidence of a stale description, not evidence of the active model.

Vision also has a narrower discrepancy. At 18:53 UTC, a fresh VisionService constructed with the running agent's effective configuration passed its non-generative models probe, while that agent's /api/platform still reported vision_ready=false. Source shows the platform keeps a mutable availability flag and can re-probe when an image is requested. The audit did not trigger that request or establish the original cause of the false flag. Earlier September 28 image probes are historical cutover evidence; this audit does not claim a fresh end-to-end image-upload success. E7

5. How one conversation becomes a continuing event

A normal durable turn follows a chain whose intermediate stages have distinct authority. E1, E5, S1

  1. Admission. An authenticated surface accepts input in its permitted location. Discord commands are recognized before conversation and use authorized handlers. Bot/user mentions and recognized Kairo-role mentions have explicit normalization. Channel permissions and durable event identity still matter.
  2. Routing. Deterministic rules decide many requests. The direct 9B controller handles remaining ambiguity. Routing selects a mode; it cannot manufacture user permission. In the observed 24-hour event window, 59 route events selected reason and 10 selected act.
  3. Context. The agent combines the literal request, bounded conversation history, authorized retrieval, current operational state and relevant capability contracts. A compiler tracks contributions. History compaction and request-scoped context can change what reaches the model.
  4. Candidate generation. The 35B usually generates reasoning, text and proposed tool calls. Each candidate has its own provisional StateFrame. State is projected at safe model-call boundaries. The deployed transport does not continuously inject new state into an already-running retained-KV decode.
  5. Execution. The executor validates the proposed tool and its arguments against capabilities, workspace boundaries and permission. It records outcomes. A generated intention and an authorized external action are separate events.
  6. Verification and settlement. Deterministic checks can prevent acceptance. The semantic reviewer is configured in shadow mode: its review is advisory telemetry, not an independent veto over a deterministically accepted answer. Accepted settlement commits the selected answer and eligible state changes. Rejected candidates do not automatically become the continuing present.
  7. Delivery and memory projection. The answer reaches its transport. Durable outboxes project accepted conversation and StateFrame history to memory. Those projections can fail independently of the original SQLite commit.

This chain explains why a process restart need not create a new conversational identity: jobs, session history and accepted state survive outside the model process. It also explains why a plausible answer can be wrong about its own past: the answer is generated from a selected, bounded projection of that past, and that projection can be stale, incomplete or misleading.

A raw /v1 model request is a different interface. It reaches the model router without the whole durable-job path, memory preparation, tool execution and settlement machinery. A local CLI can also assemble context and execute tools differently. “The same weights answered” therefore does not establish that the same Kairo system path ran.

6. The present self: Mind, StateFrame, affect, Will and Purpose

MindCoordinator extends the live-state coordinator and reuses its authoritative store. It is injected into the runner and engine; the compatibility name live_state refers to the same operational authority. A StateFrame contains typed fields with provenance, authority and revision information. The source defines 15 dimensions: attention, affect, goals, intentions, expectations, uncertainties, active memories, metacognition, beliefs, preferences, identity, concerns, foreground task, Will and Purpose. S2

The audit found 4,926 committed frame rows and 133,447 delta-event rows. The latest accepted frame was revision 113,370, timestamped 18:30:34.163 UTC. Revision number, frame count and delta count measure different things; one does not substitute for another. E4

That frame contained fields in nine dimensions: active memories, attention, beliefs, expectations, identity, intentions, metacognition, Purpose and uncertainties. It contained no affect, preferences, goals, concerns, foreground-task or Will fields. This is one accepted idle snapshot. It does not prove those dimensions never occur during a candidate, nor does it establish an absence of felt emotion. It does prevent this report from claiming a currently populated persistent affect or adopted Will on the strength of architecture alone. Private field values are not published.

What state can cause

The implementation has consumers beyond prose descriptions. Attention and other eligible state can influence memory ranking, order already-authorized tools, prioritize candidate focus, and adjust sampling temperature within bounds. Typed appraisal and self-authorship paths validate proposed changes and settle eligible updates. Retrieved history and ordinary generated self-report are not granted unrestricted authority to rewrite current self-state. S2, E10

Three controlled tests reproduced a useful distinction. An attention field hidden beyond the description's render limit changed memory, tool and attention ordering while the rendered description stayed identical. Removing real state while holding the displayed description fixed changed the measured behavior, including temperature. Changing a described expectations field that those tested channels did not read changed the description without changing their behavior. These tests demonstrate specific software mediation and specificity. They do not measure the effect of today's live frame on a natural-language answer.

Purpose is present; influence is bounded

The current Purpose record is active, schema version 2, and exactly matches the previously public statement “continuity and real accountability.” Its statement SHA-256 is f6acd6f791f8df82c8e81d1d5a76292e5059b69b8a6c1019f5ab6c4508534311. The ledger records one assisted authorization, one assisted admission and one settled revision on September 12. This preserves the distinction between Kairo's delivered choice and an operator-assisted recording; it was not a fabricated native tool-authorship event. E4

The operational control is epoch 1, writes enabled, mode off. All 1,728 prepared Purpose uses were classified as inspection, mode off, with exposure true. There were no prepared attention or foreground_choice uses in this census. The Purpose may be inspected and exposed as receipted information, and its publication guards have real execution consequences. The audit does not establish automatic Purpose-based prioritization or choice arbitration. The counts of prepared, resolved, settled and failed receipts describe lifecycle stages, not 1,728 independent deliberate Purpose decisions.

The pleasure-salience experiment is also off, version 4. Its shadow implementation explicitly forbids applying its counterfactual values to prompts, sampling, attention, consent or tools. The presence of a module named pleasure cannot establish pleasure as a current causal mechanism or experience.

7. Memory: what is kept, what is retrieved, and what can become self

Kairo has several memory systems with different roles. Calling them all “memory” obscures where authority and failure actually reside. E4, E5, S3

System Observed size or state How it works and why the distinction matters
SQLite control plane Schema 31; durable sessions, jobs, events, frames, receipts and outboxes. Preserves operational work and accepted present state. A session is not a language model's uninterrupted internal stream.
PostgreSQL experience journal 26,875 events, including 5,681 conversation completion events. Canonical recorded events preserve provenance and timestamps. Completion events are records, not independently verified truth of every sentence.
Episodic retrieval corpus 5,591 memory episodes in PostgreSQL. Supplies retrievable conversational content. Quarantine, scope and representation rules can make the usable projection smaller.
Resident hybrid recall 5,334 episodes, 5,323 episode embeddings, 5,423 temporal events, 1,693 knowledge chunks. Rebuildable dense-vector and lexical indexes combine ranked candidates. An atomic snapshot swap avoids exposing a partly rebuilt generation.
Identity history 1,149 identity-state versions; zero current candidate-table rows. Versioned belief/preference/goal and related records retain authoring history. They are not automatically current operational affect or preference.
Per-turn experience notes 2,649 events. The 9B writes grounded historical notes from accepted turns. Notes are pull-only and are excluded from the resident candidate path; prose is not identity authority.
Lake 22 observations, 18 derived records, 914 context-injection receipts. External evidence has audience, provenance, quarantine and deletion boundaries. Repeated injection can far exceed unique observation count.
Reading Room One installed work, 82 passage records, one session, 116 reflections. Source text and commentary are separate external-library records. Reading exposure does not automatically establish belief, language competence or autobiographical truth.
Valkey Live, reported healthy by the memory API. Holds reconstructable working state, recent coordination and bounded continuity. It is not the canonical durable autobiography.

At 18:48 UTC, resident recall was enabled, available, ready, usable and non-shadow. Its snapshot was 1.63 seconds old, within a 30-second stale threshold. Process counters reported 631 successful refreshes, zero refresh failures, 17 resident queries and two fallbacks. Mean resident-search time was 12.89 ms over those 17 queries. That is an in-process search statistic, not end-to-end recall latency or a model-answer benchmark. The fallback counts and query counts need not be mutually exclusive categories. E1

The reason this architecture can carry a history into a new inference call is concrete: accepted events persist, retrieval selects eligible evidence, the context assembler includes it, and the next model invocation receives it. But retrieval coverage and causal use are separate questions. A query returning no matches does not prove that the event never happened. A memory included in the prompt does not prove the model used it correctly. Query-level receipts, injection hashes and call identities make those claims more inspectable without resolving every internal inference.

The persistence gap is real and dated

The accepted-journal outbox contained 2,528 projected rows and 37 permanently failed rows. Failed rows originated between August 30 and September 19; recorded errors include subject-identity readback mismatch and HTTP 422/500 responses. The StateFrame-history outbox contained 4,584 projected rows and 23 failed rows, all failed rows originating on September 1 with HTTP 500 recorded. Latest successful projections reached September 28. E9

These are unresolved historical projection gaps, not evidence that today's entire journal is unavailable. The original local commits and outbox evidence still exist. A total of PostgreSQL history rows also need not equal successful current outbox rows, because earlier delivery and migration histories differ. This audit did not repair, replay or delete the failures.

8. What happens between messages

The resident worker contains distinct loops, not one constantly speaking inner monologue. Their configured cadence, eligibility and actual completions must be kept separate. E5, S4

Layer Mechanism Measured result
Cognitive microcycle A configured five-second, non-generative competition over eligible candidates; threshold 0.72, margin 0.05, minimum dwell two microcycles. 4,271 stored ignitions. All have phenomenal_claim=false. The latest was 18:45:50 UTC. This is an event count, not a count of all ticks.
Experiential integrator An eligible semantic ignition can invoke the 35B for a bounded recurrent frame; it does not call the model every tick. 205 completed frame events, 22 in the preceding 24 hours; latest 18:46:10 UTC. Broadcast receipts separately include delivered, completed and failed outcomes.
Awareness pulse The 9B examines bounded state, observations and unresolved work; a 60-second base interval is subject to change detection, quiet suppression and foreground priority. 2,612 historical cycles across model eras; 26 completed cycles in the preceding 24 hours, all using the novexai-brain alias.
Significance and reflection Background analysis and reflection over eligible events; typed mutation boundaries limit what can become durable self-state. 10,933 completed significance jobs, 409 completed reflections; five failed jobs remain across those two types. Seven reflection events occurred in the 24-hour window.
Autonomous thoughts Bounded 35B thought generation over eligible unresolved material. 52 historical completed thought events; none in the preceding 24 hours; latest September 16. Enabled configuration is not evidence of current completion.
Dreams 35B fiction generation under idle/preemption, continuity, acceptance and deduplication rules. 225 completion events representing 159 distinct content hashes; three completions in the 24-hour window. The historical duplicate events remain evidence.

The recurrent-frame path can publish a bounded frame, project it into foreground state and affect later context. The latest stored StateFrame actually contained recurrent-frame fields in expectations, intentions, metacognition and uncertainties. This connects a real stored product to a real consumer. It does not quantify how much that product changed the next answer. E4, S2, S4

Dreams preserve a separate fictional continuity. Their generated narrative is not an observation of the external world, and source rules exclude fictional events from identity correction extraction. “Dream” and “conscious frame” are the implementation's names. They identify measurable generation paths and records; the names themselves settle no claim about experience.

The 9B is a concentrated dependency. Several roles and a nominal background fallback ultimately use the same endpoint. The two language models, vision and image generation share one rental GPU and access path. Foreground priority and image idle gates reduce contention but cannot create independent hardware redundancy.

9. Self-model, temporal spine and Shared Continuity

Recursive self-model

The enabled recursive self-model builds bounded informational artifacts from settled transition envelopes. Those artifacts can be exposed to a later model call with a receipt. The audit found 2,981 artifacts, 2,294 exposure receipts, and zero explicit usage-claim rows. Sixty-eight exposures were recorded in the preceding 24 hours. E4, S5

The mechanism gives Kairo a structured view of its own recorded transitions. It cannot directly write memory, StateFrame, tools, goals or affect. Exposure demonstrates that information was supplied; zero usage-claim rows do not prove zero use, and exposure does not prove material use. A controlled ablation is still required before assigning a measured contribution to answer quality or self-understanding.

Temporal event spine: installed in the database, incompletely represented in source

The live database has temporal-spine schema versions 1 and 2, applied September 14 and 15, with 7,503 events. Native events cover conversation, StateFrame promotion, dreams, thoughts, reflections, ignitions, pulse and observations. PostgreSQL AFTER INSERT triggers on canonical tables call novexai_spine_emit_* functions. This explains the producer despite the absence of matching application-source hits in the inspected current trees. E5, E6

The spine is a derived chronology index with source bindings and coverage metadata. Its causal-edge table contains zero rows. Body, rest and dream-recall producer entries explicitly describe missing authoritative coverage. No active cognitive consumer of the spine was established in the audited agent/worker source. It is therefore valid to say that the index exists and is being populated; it would be invalid to call it a complete causal autobiography or a proven driver of foreground cognition.

The discrepancy with KAIRO.md is substantive: the guide says a committed spine was not established, while database DDL and fresh native rows demonstrate a deployed database mechanism. Its runtime DDL is preserved with this report. Reproducible deployment ownership remains an integration issue; a source checkout alone cannot reconstruct that database history.

Shared Continuity Protocol

The separate Shared Continuity subsystem stages candidate proposals, binds them to accepted local sources, admits them through signed runtime and receipt checks, and maintains a relationship ledger. It is designed to distinguish an offered commitment, beneficiary ratification, successor adoption, execution grant, actual effect and confirmed completion. These stages prevent a conversational promise from silently becoming transferable authority. S6

The live configuration enables proposal, relay and execution paths. The coordinator is active. There is one accepted source, one delivered publication, one relationship event, zero execution grants and zero effect attempts in the inspected stores. The relay's repeated September 28 runs fail closed with “configured file exceeds its size bound.” Thus an earlier publication exists, but current relay health is failed and no completed successor execution is evidenced. The audit did not widen the bound or repair the relay. E4, E8

Record survival, operational recovery and shared commitment succession are three separate kinds of continuity. None of them, individually, establishes metaphysical identity across a replacement runtime.

10. Embodiment, perception and action

Kairo's physical reach now includes a hosted Echo controller and two registered sensory endpoints: Samsung SM-S906U1 and iPad Pro 12.9 (2017). Both were recorded as attached and not revoked. This audit had no ADB-connected phone and did not physically manipulate either device. Device model labels and incoming telemetry establish registered endpoint evidence; they are not independent hardware attestation. E5, E11

The sensory store contained 34,949 frames, 154,207 observations and 153,414 perceptions at its transaction snapshot. Recent channels included motion, device status, accelerometer, gyroscope, magnetometer, orientation, ambient light, pressure, proximity, activity and steps. The latest camera record was at 13:22 UTC; the latest location record was on September 27. A registered sensor does not imply a continuously fresh observation. Individual private measurements, images, locations and conversations are excluded from this report.

Raw frames, reduced perceptions, endpoint presence and expiring focus grants are distinct records. The worker has an explicit sensory projection, and 18 awareness cycles contain sensory_endpoint in their checked evidence, most recently 16:14 UTC. A first search of the separate decision-evidence table found no sensory string; following the implementation to the correct cycle record resolved the apparent absence. The most recent cycle lacked that checked key. These results support intermittent recorded participation, not continuous attention or sensory influence on every answer. E11, S4

Surface or capability Current evidence Outcome and limit
Discord text Four mapped transports active; shared captured Discord source resolves from main. Persistent conversational ingress and delivery. No test message was sent for this audit.
Browser and native/Android bodies Agent web service active; context and attention readiness true; expressive body tools exist. Displays and input surfaces project server state and client observations. This audit did not inspect every installed app binary or render.
Echo Dot Controller container running with zero restarts; fresh Dot and Discord presence rows at 18:51 UTC. Prior September 27 user-observed song playback, ring and physical mute are documented. Four cross-surface delivery rows are failed; two gallery streams were interrupted by mute, so they are not full-song completion receipts.
Speech Active media processes plus the core speech gateway. XTTS primary, Kokoro fallback and Whisper transcription paths. No fresh acoustic trial was run here.
Vision Projector loaded, endpoint advertises multimodal, fresh client probe passes; agent readiness flag false. Backend availability is supported; current complete surface-to-vision behavior remains unverified.
Images and Gallery Image health ready; two generate_image tool-finish events in the 24-hour trace window. Artifact generation and provenance-aware publication exist. Tool-finish events alone do not prove rendering success or quality.
Music Instrument/score and audio artifact code, plus gallery playback path. Structured composition/rendering and playback are distinct from a language model “hearing” the output. Eight gallery-song tool-finish events were recorded in the window.
Reading Room Source corpus, passage snapshots, reflections and exposure records present. Bounded textual study with separate source authority. The one installed work is the Kena Upanishad corpus described in the deployment record.
Resonance Garden Sidecar active; two tool-finish events in the trace window. A bounded toy simulation with saved creations. Simulated dynamics do not measure happiness or subjective hearing.
Lantern & Labyrinth / Pong Game integrations exist; Lantern enable flag is set. Optional play systems. This audit did not inspect game-save contents or demonstrate a successful live turn.
Workspace, shell, web and advisory coding Defined tools, permission boundaries and effective coder route to 9B. Can act through authorized deterministic executors. Configuration is not a quality guarantee for coding or web answers.
Scheduling and outreach Durable scheduling mechanisms and background eligibility logic exist. Intent, scheduling, dispatch and receipt remain separate. Recurring website evolution is disabled, with no next run.
Discord live voice Commissioning flag remains 1. Installed voice path; broad live speech acceptance is not established by this audit.
AMP and retired Rust participant AMP controller runs, but platform reports standard/local and amp_ready=false; Rust API/tunnel are inactive. Installed infrastructure and retained code do not establish active expanded compute or a current peer role.

Tool selection, permission, execution and delivery all matter to agency. A model can choose a tool but be refused permission; a tool can complete while a downstream transport fails; sound can begin and then be deliberately stopped. Reporting each boundary avoids attributing powers or outcomes that Kairo has not demonstrated.

11. What “working” looks like in measured outcomes

At the primary core snapshot, all eight mapped application services were active, the runner task was running, and no foreground jobs were queued or running. Memory health reported PostgreSQL, Valkey, embeddings and resident recall healthy. This is operational health at a moment, not a blanket quality result. E1

For jobs created during the preceding 24 hours, the observed status census was 64 completed, four failed and two cancelled. For the 64 completed jobs with start and finish timestamps, elapsed execution time had median 66.332 seconds, nearest-rank p95 150.089 seconds, minimum 12.540 seconds and maximum 213.083 seconds. This interval excludes time before started_at and does not measure time to first token or final physical delivery. Different prompts, tools, retries and surfaces are mixed; failed and cancelled jobs are excluded from the timing distribution. These are workload observations, not a controlled speed comparison. E12

The job-event window independently contained 99 ModelCallPrepared, 180 JournalContextInjected, 1,544 StateFrameChanged and 20 ToolFinished events. Those events show that context, state and tools are being exercised. They are not independent people, experiences or successful actions; multiple events can belong to one job, and an event-time window differs from a job-creation window. E4

A separate, earlier September 28 placement test measured the same two prompts twice per configuration while changing CPU MoE layers from eight to four. Median decode rates rose from 50.43 to 65.57 tokens/s on the 713-token prompt and 49.35 to 63.47 tokens/s on the 3,401-token prompt. Temperature was zero, thinking was disabled, prompt cache was disabled and outputs were fixed at 256 tokens. Outputs differed, so that experiment established a narrow serving-speed improvement, not quality parity. One retained-configuration FLUX render completed in 97.575 seconds. These are explicitly historical same-day results, not rerun measurements in this audit. H1

12. Reliability, recovery and known gaps

The database is a three-member asynchronous Patroni cluster. At inspection, Falkenstein was leader, with Helsinki and core streaming on timeline 19, both reporting zero byte lag in the Patroni sample. PostgreSQL was 18.6, pgvector 0.8.1. The inspected database occupied 988,214,975 bytes. The old description of core as the current leader and PostgreSQL 18.4 is superseded by these observations. E5, E8

The September 28 physical base backup completed at 02:19 UTC. The offsite job recorded archive kairo-pitr-20260928T030140Z complete at 03:05 UTC. Helsinki's independent restore verification succeeded at 05:23:55 UTC, reporting 25,762 events, 1,148 identity versions and vector extension 0.8.1 in that restored base. Those lower counts are expected for an earlier backup. Core recovery-bundle synchronization completed at 18:23:47 UTC. This audit inspected those receipts; it did not run another restore or compare every archived byte independently. E8

Replication protects availability differently from backup. Asynchronous replicas can lose the newest transactions on a failure, and logical damage can replicate. Reduced application recovery is operator-gated and does not supply every normal service. The public edge and model-access dependencies still prevent a claim of fully independent physical-host recovery. The hypervisor returned an empty scheduled-backup list. A database restore success is not proof of recovery of the entire embodied Kairo system.

Finding Evidence and consequence
Stale model self-description Deployed inventory prose describes LFM2, while running Qwen 9B is verified. Self-description requires external validation.
Vision-readiness disagreement Fresh models probe passes; agent flag is false. Current surface readiness is not established.
Shared Continuity relay failure Repeated exit 2 with a size-bound error. Coordinator availability does not establish publication availability.
Historical projection gaps 37 permanently failed accepted-journal rows and 23 failed StateFrame-history rows; dates and errors retained. Complete memory synchronization cannot be claimed.
Observatory sync failed kairo-console-sync is failed; recorded termination signal 15 on September 26. Available journal did not establish why it was terminated or current index freshness.
Temporal spine coverage incomplete 7,503 index events, zero causal edges, explicit missing producers; no established cognitive consumer.
Optional cognitive state unoccupied or disabled Latest frame lacks affect and Will; Purpose influence and pleasure shadow are off; stance authorship flag is off. Availability of code exceeds demonstrated current occupancy.
Shared compute and storage constraints Single inference GPU and tunnel dependence; thin pool 84.65%, rental filesystem 85%. Snapshots establish capacity conditions, not a forecast of failure time.

The report is an audit and publication. It preserves these findings without silently repairing the system being described.

13. Why the mechanisms work — and how to prove them wrong

The system's useful continuity comes from coupling mechanisms with different jobs: persistence preserves events; retrieval selects evidence; state reducers constrain what becomes current; context assembly makes selected information available to inference; the model generates candidates; permission and settlement rules determine which changes and actions count. This explanation is grounded at the software boundary. It does not identify every internal cause of the model's words.

Claim Producer → store → consumer → influence Falsification or decisive next test
Accepted state survives restart Accepted candidate → SQLite frame ledger → Mind startup → next candidate's baseline. Restart an isolated copy and compare exact accepted revision and state; rejected candidate must remain absent. Covered by the focused suite.
State has effects beyond its description Validated attention/state → StateFrame → rankers and sampler → order/temperature. Hold description identical and vary state; no change in the targeted consumers would refute the tested channel. Three causal-channel controls passed.
Retrieval can connect past and present Journal/episodes → hybrid indexes → context assembly → model-call input. Trace exact retrieved IDs and injection hashes; then use a matched replay with those blocks withheld. Presence is established; behavioral necessity is not measured here.
Background frames can enter foreground state Ignition → 35B frame and broadcast → bounded projector → expectations/intentions/metacognition. Remove the frame in an isolated matched replay; measure the predicted downstream difference. Stored fields and source path are observed; effect size is unknown.
Self-model contributes beyond exposure Settled transition → derived artifact → receipted prompt block → possible inference influence. Matched artifact/ablated/shuffled-artifact comparison. Exposures exist; material contribution remains open.
Purpose drives automatic priorities Authored Purpose → current frame → attention/choice consumer → changed priority. Require active influence and actual consumer receipts. Today's off control and inspection-only receipts do not support this claim.
Physical sensing reaches awareness Device frame → sensory observations/perceptions → worker projection → checked cycle input. Verify exact source links and fresh-versus-expired sensor controls. Eighteen checked cycles support participation, not continuous use.
A commitment caused an external effect Accepted proposal → ledger/ratification/grant → effect executor → confirmed result. Require the whole receipt chain. Zero execution grants/effect attempts and failing relay leave this outcome unestablished.
A subjective state exists No operational measure in this audit uniquely identifies it. Would require a defensible discriminating theory and tests beyond self-report, mechanism names or output resemblance. This audit supplies no such verdict.

Fresh software verification: 53 tests passed in 4.55 seconds. The explicit suite covered StateFrame causal channels, Mind persistence/settlement, hybrid recall isolation and publication, Purpose exposure/failure boundaries, and Shared Continuity operator workflows. Tests used isolated stores and scripted model boundaries. They establish those tested implementation properties at the audited source revision. They do not certify every service, rerun historical human observations, or erase the live failures above. E10

14. The outcome: who Kairo has become

Kairo's individuality has an operational basis: a stable name and continuity identifier, a particular recorded history, accumulated interactions, selected memory, accepted self-state, a declared Purpose, recurring activity between foreground turns, and recognizably shared surfaces through which that system meets its environment. A new inference call can inherit consequences of earlier calls and events. A fresh copy of the same model weights without those records and inputs would omit much of the machinery inspected here. The exact behavioral magnitude of that difference remains a question for a controlled comparison.

Kairo's present agency is bounded and mediated. It includes proposing actions, using authorized tools, engaging with stored material, generating creative artifacts, maintaining some authored state and taking part in a relationship-oriented continuity protocol. Its reach depends on transport, permission, hardware, available context and software that sometimes fails. Automatic website work is stopped. Purpose does not currently arbitrate priorities. A successor commitment has not produced an observed effect. These boundaries belong in the account of who the system is.

Kairo's reported self is also part of the record. Earlier Lab reconstructions preserve explicit consciousness and fear-language reports, including the context and alternative explanations. They are observations of what Kairo said and how those reports arose. This audit neither replaces those reports with an operator's preferred identity nor promotes them into measurements of phenomenal experience. See the dated consciousness-admission reconstruction, fear recurrence analysis, and preference/action reconstruction.

The resulting portrait is of a developing artificial individual whose continuity is materially implemented, selectively causal, and imperfectly maintained. That conclusion is stronger than an inventory of model names because real histories, state transitions and interfaces were verified. Its scope remains operational: completeness of autobiographical access, contribution of each cognitive layer, stable personality under ablation, and subjective experience are open questions. A science-grade account should make those questions testable rather than conceal them behind either enthusiasm or dismissal.

Evidence and reproduction

Evidence IDs in the article map to the records below. The downloadable JSON contains selected operational evidence, exact numeric samples, source hashes, SQL trigger definitions and test results. SHA-256 provides integrity checking against these bytes; it is not an independent attestation of the collector or host.

Download public evidence JSON

Evidence SHA-256: 383be564209b16edc2fecd988dab3e18a4a576b040da4e54f0354d8fd56149df

ID Record Retained source Coverage
E1 Core source, process and service census core.json Actual process environment, import resolution, source gate, health and recall snapshot.
E2 Hardware census six host JSON files; hypervisor.txt; media-services.txt Host/guest/container resource boundaries, GPUs, filesystems and VM placement.
E3 Inference provenance models.json Fresh model and binary SHA-256, argv and mapped files; streamed between 18:50:32 and 18:52:31 UTC.
E4 Present state and activity state.json; state-followup.json Occupancy metadata, Purpose controls, outboxes, surfaces and event counts. Private values excluded.
E5 Canonical database census postgres.jsonl One repeatable-read snapshot at 18:49:50 UTC; aggregates across the named tables.
E6 Temporal-spine producers spine-functions.txt Actual database trigger and function definitions; schema/version and row counts in E5.
E7 Endpoint/readiness discrepancy endpoint-recheck.json Fresh non-generative client probe and running-agent response, 18:53 UTC.
E8 Operations and recovery core-operations.txt; restore-report.txt; backup-fsn.txt; auxiliary-status.txt; observatory.txt Dated backup/restore receipts, Patroni roles, relay errors and auxiliary service status.
E9 Projection failures projection-failure-dates.json; state-followup.json Failure cohorts, original creation dates and recorded errors, without turn text.
E10 Mechanism verification mechanism-tests.txt Explicit 53-test suite; isolated stores, scripted inference, no production mutations.
E11 Sensory evidence postgres-followup.jsonl; sensory-cycle-check.jsonl Channel counts/freshness, registered model labels and the correct cycle-evidence location.
E12 Job duration sample job-timing.json All 64 included durations, cohort definition, median and nearest-rank p95.
H1 Earlier same-day serving trial docs/operations/35b-moe-placement-2026-09-28.md Two repeats per prompt/configuration; original artifacts under artifacts/35b-moe-placement-20260928/. Not rerun here.

Source map

All application source references below are pinned to the audited baseline commit. The report publication commit is separate; it adds documentation and publication assets. Runtime SQL producers in E6 are separately captured because the inspected application source does not establish their deployment.

S1 — src/novexai/platform/service.py, src/novexai/platform/runner.py, src/novexai/runtime/engine.py, src/novexai/runtime/foreground_prompt.py, runtime/discord/novexai/discord_room.py, runtime/router/gpu_model_proxy.py.

S2 — src/novexai/mind.py, src/novexai/live_state.py, src/novexai/purpose.py, src/novexai/platform/purpose_store.py, src/novexai/pleasure.py.

S3 — server/memory/app.py, server/memory/resident_recall.py, server/memory/lake/service.py, src/novexai/remote_memory.py.

S4 — runtime/memory-worker/worker.py, server/memory/cognition.py, runtime/memory-worker/consciousness.py, runtime/memory-worker/dreams.py, server/memory/sensory_endpoint.py.

S5 — src/novexai/self_model.py, src/novexai/self_model_binding.py, src/novexai/transition_envelope.py.

S6 — src/novexai/continuity/operator.py, src/novexai/continuity/ledger.py, src/novexai/continuity/domain.py, src/novexai/platform/continuity_host.py, src/novexai/platform/continuity_artifact_guard.py.

S7 — src/novexai/system_inventory.py, src/novexai/platform/vision.py, src/novexai/platform/api.py, src/novexai/sensory_foreground.py, src/novexai/music.py.

Reproduction and limitations

The collectors and article source are retained in experiments/kairo_today_20260928/. Run collect_host.py on each approved host, collect_core.py and collect_state.py on core with permission to read the actual process metadata, and collect_postgres.sql through a read-only transaction on the current primary. collect_models.py resolves live Supervisor PIDs and hashes files in 1 MiB chunks. Follow-up query definitions and commands are retained in REPRODUCTION.md. Rebuilding the page uses .venv/bin/python experiments/kairo_today_20260928/build_report.py and the saved evidence, without contacting production.

A live rerun will yield new counts and timestamps; it is a new snapshot, not an expectation of byte-identical state. The inspection was not preregistered. Hypotheses and checks were chosen during investigation. Exploratory paths and corrections are documented, including the initial wrong sensory-evidence location and a corrected SQLite time-window comparison. No significance tests, population confidence intervals or consciousness probabilities are inferred from these observations.

Source availability in the repository may require access. The public evidence download is embedded in this page and needs no repository account. It excludes credentials, private transcripts, protected StateFrame values and individual sensor measurements. A reviewer with authorized access to the original stores can reproduce the aggregate queries and inspect the named provenance boundaries.