When we say we track engagement, here is exactly what we measure. Every signal is derived from the viewer's camera feed using on-device AI — no video ever leaves their device.
AttentionTag's AI models run entirely on the viewer's device (edge computing). The models analyse the camera feed locally and produce lightweight engagement signals — small data points like "focused" or "positive valence" — which are sent to the server. No images or video are ever transmitted. This privacy-first architecture means we can give you rich engagement insights without compromising anyone's visual data.
Understanding how your audience feels
The base emotional expression detected on the viewer's face at any given moment.
The overall emotional tone — is the viewer in a positive, negative, or unclear emotional state? Derived from the mix of detected emotions.
Also tracked as a continuous 0–1 score for finer granularity.
The strength of emotional arousal — how strongly the viewer is reacting, regardless of whether the emotion is positive or negative.
Also tracked as a continuous score for analytics.
Why three signals for mood? Emotion gives you the specific expression. Valence simplifies it into positive vs. negative — useful for quick dashboards. Intensity tells you the strength of the reaction — a mildly happy student and an extremely happy student both have positive valence, but very different intensity. Together, they give speakers a nuanced picture of how content is landing.
Is your audience paying attention?
Uses gaze estimation to determine whether the viewer is looking at their screen or looking away.
Detects whether the meeting or learning application is the active window on the viewer's device, or if they've switched to another app.
A combined signal: the viewer is considered truly focused only when both their gaze is on-screen and the meeting app is active. This filters out false positives from either signal alone.
Are viewers awake and present?
Monitors the eye aspect ratio to detect when a viewer's eyes are drooping or closed for extended periods — an early indicator of drowsiness or fatigue.
Detects whether a face is visible in the camera frame. This tells you if the viewer is physically present at their device or has stepped away.
Tracks whether the viewer's camera is turned on or off. When the camera is off, all vision-based signals are paused — we simply record that the camera was inactive.
All signals in one view
| Signal | Category | Possible Values | What It Tells You |
|---|---|---|---|
| Emotion | Emotion & Mood | Happy, Neutral, Sad, Angry, Fearful, Surprised, Disgusted, Contempt | The specific facial expression detected |
| Valence | Emotion & Mood | Positive, Negative, Unclear | Overall emotional tone (positive vs. negative) |
| Intensity | Emotion & Mood | Intense, Neutral | Strength of emotional arousal |
| On-Screen Focus | Attention | Focused, Distracted | Whether gaze is directed at the screen |
| App Activity | Attention | Active, Inactive | Whether the meeting app has window focus |
| Effective Focus | Attention | Focused, Distracted | Combined gaze + app activity (true focus) |
| Drowsiness | Alertness | Awake, Sleepy | Fatigue detection via eye aspect ratio |
| Presence | Alertness | Present, Absent | Whether a face is visible in the frame |
| Camera Status | Technical | On, Off | Whether the viewer's camera is active |
All vision models run on the viewer's device. Only lightweight signal labels (like "focused" or "happy") are transmitted — never images or video. This edge-computing architecture ensures privacy while delivering rich engagement analytics.
Read our full Privacy Policy for details.