The AI that reads the room.
Face, eyes, hands and voice — now with posture, the objects in the room and everyone else in frame.
Every reading still traceable to the movement behind it.
standby
00:00
Camera— fps
One permission covers both. Video and audio stay on your device — nothing is uploaded, nothing is recorded to disk.
🙂configuration —
profile 0%
Plus modulesnot loaded
Three models run beside the face: a body landmarker for posture and gesture, an object detector for
what is in the room, and up to four faces at once. Everything stays on your device. Turn a module off if the frame rate drops.
Measurement quality—
—conditions, not correctness
Waiting for a face.
Below 50% no configuration is declared at all. This score says whether the conditions allow a measurement —
it says nothing about whether the resulting interpretation is right.
Action units
No muscle is moving appreciably above your resting face.
Hands on face
Hands not detected yet.
Events
Every muscle and every hand contact is tracked from onset to offset. Anything under half a second is logged as a flash.
Last minute
60 s agonow
Waiting for a face
Sit in front of the camera with the light on your face, not behind you.
From EMFACS rules
From your photos
No reference photos yet. See the Photos tab.
upper face—
lower face—
onset—
held—
stability—
symmetry—
head—
gaze—
blinks—
voiceoff
body—
in frame—
Where you look, relative to the camera. This is a proxy for looking at a person: it measures gaze toward the lens, so put the other person's window near the camera if you are on a call.
at camera—
longest hold—
breaks—
while speaking—
while listening—
dominance ratio—
2 min agonow
Blink dynamics—
rate—
mean interval—
longest gap—
blink duration—
bursts—
your baselinelearning
What moves blink rate. Screen use and reading suppress it, often to a third of resting rate. Dry air, contact lenses, tiredness,
antihistamines and general arousal push it up. Parkinson's lowers it markedly; some tic disorders raise it. It is a real physiological signal with many causes —
and it is not a deception cue: meta-analyses find no consistent relationship, in either direction.
Reference values, for comparison only. Argyle's classic work put eye contact in Western conversation at roughly 61% overall — about 41% while speaking and 75% while listening.
A 2018 dual eye-tracking study found people look away from their partner's face about 29% of the time while speaking and 10% while listening,
with mutual face gaze around 60% of a conversation in short bursts averaging 2.2 seconds. Binetti et al. found the preferred length of a mutual gaze is about 3.3 seconds, comfortable between 2 and 5.
The visual dominance ratio divides gaze-while-speaking by gaze-while-listening: around 0.6 is typical, near 1.0 is associated with dominance.
Norms are Western and vary a lot by culture and by person — direct gaze is read very differently in parts of East Asia, and autistic people often find it costly. Do not treat a low number as a fault.
Tiredness index—
—
alertmarked
PERCLOS—
yawns0
long closures0
over 2 s0
blink duration—
lid droop—
confidence—
What each measure is.PERCLOS is the share of time the eyelid covers most of the pupil, measured over a rolling minute;
it is the single best validated drowsiness measure and the one used in vehicle monitoring. Under 6% is the alert range. Long closures are lid closures
between half a second and two seconds — too slow for a blink. Anything over two seconds is counted separately because that is the microsleep range.
Blink duration lengthens with time on task before the rate itself moves. Lid droop compares eye aperture in the last third of the session
against the first third, so it is a drift against you, not against a population.
How yawns are counted. The jaw must open wide and stay open for at least 1.2 seconds while the eyes narrow or close, with no
voice at the same moment. The hold is what separates a yawn from speech; the eye closure is what separates it from a laugh or a wide-open surprise.
A covered mouth defeats it entirely, so treat the count as a floor rather than a total.
Would look the same if. Dry air, contact lenses, air conditioning, allergies, antihistamines, a screen at the wrong distance, a bright
window behind the camera, or simply a long stretch of reading. These measures index ocular fatigue, which overlaps with sleepiness but is not the
same thing, and none of them can tell you why. High readings here are worth acting on; they are not worth diagnosing.
Stored sessionsnone yet
sessions0
total time—
mean tiredness—
mean PERCLOS—
mean blink rate—
mean gaze at camera—
most common state—
Trendtiredness index
—latest
Sessions
No sessions stored yet. Run a session and press End & report — it is saved here automatically.
Calibration0 labelled
Export
JSON keeps every field and can be re-imported here. CSV is one row per session with the headline
numbers, for a spreadsheet or a notebook. Both export exactly what the filter above is showing.
Session review
Push to a serveroff
Optional. Leave it blank and nothing ever leaves the browser. If you fill it in, each session is POSTed
as one JSON record and marked as sent. animo-server.mjs is a working receiver you can run with no install.
Where this data lives. Sessions are held in this browser's IndexedDB, on this machine, under this origin. They survive
reloads but not a cleared browser store, a different browser or a private window — so export anything you want to keep.
Before you point this at a server. These records describe a person's face, gaze and eyelids over time. While they stay on your
own machine and unnamed, they are ordinary personal data. They become special-category biometric data under UK GDPR once they are used to
identify or single out an individual — which is what a per-person history is for. If sessions from other people will land in that database you
need their explicit consent, a retention limit, and a way to delete a person on request. The label field is free text so you can use a
code rather than a name.
input
sensitivity1.0×
pitch
—
loudness
—
rate
—
range
—
pauses0
pause rate—
filled pauses0
filler rate—
pitch range—
median pitch—
phrase endings—
jitter—
A filled pause is a held sound with flat pitch and steady energy — “mmm”, “ehh”, “uhh”.
It is found acoustically, not by recognising words: nothing you say is transcribed, and no audio is stored.
activation
—
The microphone is off. Your voice profile builds itself over the first twenty seconds you speak; after that every reading is a departure from your own habitual delivery.
What voice gives you and what it does not. Acoustics recover activation well — how fired up or flat someone is.
They recover the sign of an emotion badly: excitement and anger have nearly identical prosody, high and loud and fast. So this panel reports an activation index and refuses to print an emotion label.
Nothing you say is transcribed or recognised: only pitch, loudness, rhythm and pauses.
Spontaneous facial self-touch is a regulatory behaviour, not a message. What carries information is
how often it happens against your own baseline, where it lands, and how long it lasts — brief contacts and sustained resting are different things.
onsets—
rate—
vs baselinelearning
brief (<2 s)—
sustained—
T-zone share—
Where the hand lands
What each region is associated with
The evidence, precisely. A Leipzig group has studied this directly. Grunwald et al. (2014), replicated by Spille et al. (2022) on sixty participants,
found cortical theta power drops just before a spontaneous face touch and rises right after — the neurophysiological event is the onset, the reach and first contact,
not the sustained contact. Mueller, Martin and Grunwald (2019) found that contact duration and point of touch both shift with cognitive and emotional load.
People do this 400–800 times a day, mostly without noticing, and when Spille et al. (2022) stopped participants from doing it, memory performance dropped in the frequent touchers.
Instructed touching produces none of these effects: spontaneity is the whole thing.
So the location does carry information — about load and self-regulation. It does not identify an emotion, and it says nothing about honesty.
Waiting for a body
Step back until your shoulders and hands are inside the frame. The body model needs a torso, not just a head.
posture—
openness—
arms—
lean—
shoulders—
torso—
gesture zone—
movement—
hands visible—
vs baselinelearning
Postural events
Crossing your arms, leaning in, turning away, raising your shoulders and dropping your hands out of
gesture space are each tracked from onset to offset, the same way muscles and hand contacts are.
What posture gives you and what it does not. Body movement carries activation and engagement reliably: how much someone
is moving, how much space they take, whether they orient toward you. It carries specific emotion badly, and it carries
honesty not at all. The old catalogue that reads crossed arms as defensiveness has never replicated: people cross their arms because
they are cold, because the chair has no armrests, because it is comfortable. What is informative here is the shift against your own
baseline over the session, and the coincidence between a postural change and something that happened in the conversation.
Two stages run on the same frame. A detector finds where things are and names them from 80
everyday COCO categories — reliable, coarse. Each box is then cropped, enlarged and passed to a 1000-category ImageNet
classifier that says what it is more precisely: "cup" becomes "coffee mug" or "water bottle". Stage one is used for what is in the
room, what is in your hand and what is near your face; stage two only ever refines a name.
in frame now—
distinct seen—
people detected—
screen or phone—
near the face—
in a hand—
30%
Lower the sensitivity and the detector will name more things and be wrong more often;
raise it and it only speaks when sure. It changes nothing that was already recorded — only what is admitted from here on.
Right now
Nothing recognised yet. The detector only knows 80 everyday categories, and it needs the object
reasonably large and well lit.
Seen since the module started
Fine labels
Nothing named finely yet. This second pass crops each detected box, enlarges it and asks a
1000-category classifier what it is — so it needs the detector to find something first.
Read this as context, not as content. A phone in frame changes how to read a low blink rate and a dropped gaze —
both are what screen use does. A cup near the face explains a hand contact that would otherwise be logged as self-touch.
That is the whole job of this panel: to give the other numbers their alibis. Names come from the COCO category list and are
often wrong on unusual objects.
Up to four faces are measured at once. One of them is the subject — the largest, tracked
between frames — and only that one gets the personal baseline, the voice, the blink history and the session report.
Everyone else is read on absolute values, which is a coarser measurement.
faces now—
most at once—
subject alone—
with others—
nearest other—
mutual facing—
Each face in frame
Only one face in frame. This panel fills up when someone else enters.
Say this out loud before you point it at anyone. A second face in frame is a second person, and that person has not
consented to being measured. Nothing is uploaded and nothing is written to disk, but a reading of someone who does not know
they are being read is the exact thing this tool should not be used for. The secondary readings are also weaker:
without a resting baseline for that face, a person with a naturally downturned mouth reads as sad and a person with heavy brows
reads as angry. Treat them as a presence count, not as a verdict.
Tap a suggestion above to see which action units that label should show.
Drop photos here
Or click to choose them. One frontal face per image. Three to eight photos per label works best.
How they are used. Each photo gives 16 action-unit values; the vector is normalised and averaged per label.
Live, it compares the shape of your muscle pattern against those averages, not the faces themselves, so photos of other people still work.
What you get is your own criterion, not ground truth — eight smirks filed under “contempt” will teach it smirks.
Start a session, have your conversation, then close it: you get the percentage of time in each configuration, every muscle event counted, where you looked, hand-to-face contacts and how the voice moved.
Auto-calibrationlearning
Nothing to do. Each coefficient tracks its own floor (your resting face) and ceiling (your range), and readings become a percentage of that span. Every coefficient also has a minimum plausible range, so a muscle that never moves can never be amplified into a false signal.