Build a house that never existed, watch Lumi learn it, then judge what it would have done.
Add everyone. They share the devices and each has their own rhythm — which is the only arrangement where the household model earns its place, and the only one where two people wanting the same lamp at different hours can actually be seen. Preferences stay personal: Lumi never pools those.
Pick the devices. Each kind comes with its own abilities — a fan has three speeds and no brightness — because that is what the system reasons about.
What normally happens, and roughly when. Leave the time blank and it lands anywhere in that part of the day. Nothing fires every single day — a routine that never slips is a machine, not a person.
This wipes and rebuilds the test database on the Pi. It cannot touch real data — the server refuses to start on the real environment.
Habits are attributed to whoever the events belonged to.
DBSCAN clustered the events it was given. Confidence is the share of days the habit held.
Read from the device's own description at pairing — real ranges and step counts, not a label. A fan's three positions and a dimmer's 0–254 are different objects, and everything variable is normalised to one scale before a model ever sees it.
Two stores. Episodes vote — they are a quarter of the decision weight, and the only channel that knows whether an action was allowed to stand. Beliefs never vote — they go into the language model and change how Lumi talks, never whether a light comes on.
Trained on those events plus generated "nothing happened" samples, then calibrated so the number means what the decision assumes.
Each row is one moment it would act or ask, scored by the real decision fusion. Keep it or revert it. Your answers are recorded exactly the way a spoken yes or a switch at the wall would be — then the models relearn.
Pick a person, a device and a day, then drag through the clock. Every five-minute slot was scored through the same decision path the house runs on, so this is what Lumi would really think — not a sample of what it happened to propose.
The list above is judged by models that have already seen that day. This is the honest version: for every day, the future is deleted, the habits and models are rebuilt from what came before, and only then is the day scored. It also compares against a deliberately stupid rule — "act at the hour this device is usually used" — because a number with nothing beside it is a claim, not a finding. Takes a minute or two.
Click a day to see it in five-minute steps — blank where nothing happened.
Before your verdicts, and after. If nothing moved, the feedback did not reach anything — which is itself worth knowing.