loading
Back to the blog

[ Research ]

The face is not the meeting

Why an AI customer engineer should understand shared work, not guess emotions from a webcam.

Bhavani Kalisetty·August 5, 2026·8 min read

The face is not the meeting

Your customer looked away. Were they bored, reading logs, replying in Slack, checking a second monitor, or following the setup guide?

A human customer engineer does not know from the face alone. We use the rest of the room.

We know what question was asked. We see which screen is shared. We remember the setup that failed last week. We notice that the page did not change after the click. We hear someone say, “That is not what I see.” We ask where they are stuck. Sometimes we also notice a look—but we interpret it inside a situation, and we can still be wrong.

If we build Rover, live in the call, the goal should not be to automate the most fragile part of that judgment.

The goal should be to understand the work.

A face cannot establish intent

There are three claims that emotion-recognition products often slide between:

  1. A face moved in a detectable way.
  2. An observer would label the visible expression with an affect category.
  3. The person internally feels an emotion—or intends, understands, trusts, or agrees.

Those claims are not equivalent.

A major review by Lisa Feldman Barrett and colleagues examined whether familiar facial configurations are reliable, specific, and generalizable markers of emotion. The evidence did not support using a facial movement as a diagnostic readout of a person's internal state. People vary across situations and cultures; the same movement can appear with multiple emotions or with no emotion claim at all. The authors are equally clear that faces still communicate social information. “A face contains information” is not the same as “this face proves boredom.”

Read the review.

In a work meeting, even a perfect description of the face cannot establish the thing a customer engineer needs to know:

  • Are they following the setup?
  • Did they understand the tradeoff?
  • Does this person have authority to approve the change?
  • Is silence reflection, a muted microphone, a side conversation, or a broken connection?
  • Did they look away from Rover—or toward the log line Rover just asked them to check?

The face is one camera crop. The meeting is a coordinated task.

What multimodal research succeeds at

There is real technical progress in multimodal affect classification, and dismissing it would be as careless as overselling it.

An ACL 2023 paper built a system that combines text, audio, and the matched face sequence of the speaker. It improved emotion classification on the MELD benchmark. The paper addresses a genuine problem: in a multi-party scene, the wrong face can introduce noise into the prediction.

Read the ACL paper.

That result establishes performance on its evaluated task. It does not establish reliable access to private emotion in an enterprise call. MELD contains about 13,000 annotated utterances from 1,433 dialogues in the television series Friends. That is useful benchmark data. It is also scripted, produced, framed, and labeled differently from an engineer debugging SSO while Slack flashes on another display.

Read the MELD dataset paper.

The right scientific posture can hold both facts at once:

  • multimodal context can improve classification of expressed affect in bounded datasets;
  • that success does not validate silent inference of internal emotion, comprehension, engagement, authority, or buying intent in ordinary work.

The real meeting is multi-person, multi-device, and task-shaped

Remote-meeting research makes the webcam shortcut even weaker.

A CHI study combined large-scale Microsoft 365 telemetry with a 715-person diary study. Multitasking varied with meeting length, size, time, and type, and participants described both negative and positive forms. Answering email can be distraction. It can also be the work the meeting just created.

Read the remote-meeting study.

Now add the configurations we encounter every day:

  • one person on a laptop with the shared screen hiding the gallery;
  • a second monitor carrying logs, Slack, or the setup guide;
  • several people represented by one conference-room camera and microphone;
  • a remote participant who cannot see the whiteboard or side conversation;
  • the user sharing a customer product while Rover is active in a different authorized tab; or
  • a mobile attendee listening while someone else drives.

Hybrid-meeting research describes asymmetries in visibility, audio, resources, technology access, and power between the room and remote participants. A tile cannot tell Rover who saw the error, who controls the screen, who owns the account, or who can authorize an action.

Read the hybrid-meeting guidance.

The legal and consent line

The legal boundary is narrower and more precise than “emotion AI is banned everywhere.”

The EU AI Act prohibits emotion recognition in workplaces and education institutions except for medical or safety reasons. Other biometric, privacy, employment, recording, and wiretap rules vary by jurisdiction and use, so teams need advice for the markets where they operate.

Read the European Commission guidance.

Our product boundary should not wait for the broadest possible legal interpretation.

Rover should not create an employee or customer emotion score. It should not silently label engagement, truthfulness, confidence, or purchase intent. It should not inspect private Slack, email, unrelated tabs, or unrelated applications because they might contain “context.” Participants should know Rover is present, what it can access, what it is retaining, and how to pause or remove it.

Consent is not a modal clicked once by an administrator. It is an operating state the meeting can see.

A better context model

Rover, live in the call, should build an account-scoped model from the surfaces that make the task legible:

  • account history and the agreed success plan;
  • the current conversation and active-speaker timing;
  • meeting chat made available to the bot;
  • the shared screen or a Rover-enabled product session;
  • current product state and account-scoped logs;
  • explicit participant commands and approvals; and
  • the role and policy boundaries governing the requested action.

Diagram showing consented account history, conversation, meeting chat, shared screen, product state, logs, and explicit commands flowing into a shared-work model, which chooses guide, ask, act, verify, or escalate

Participant video may support consented mechanics such as identifying the active speaker or helping participants coordinate attention around a shared artifact. It should not become a psychological profile.

Repair the work, not the person

A good customer engineer watches for objective repair cues:

  • the same question is asked again;
  • an action fails or a visible error appears;
  • the user loops between the same pages;
  • the expected state change does not happen;
  • someone interrupts or corrects the instruction;
  • a long silence follows a concrete next step;
  • the shared screen and the spoken description disagree; or
  • an approver has not actually granted permission.

None of these proves confusion. Each justifies a small repair.

“I may be looking at a different environment. Are you in production or staging?”

“The save did not produce the expected status. May I open the account-scoped logs?”

“I can highlight the control on your shared page, or take the step in your Rover-enabled session. Which do you prefer?”

That is a safer and more useful loop than assigning a hidden score to a face.

Joint attention before autonomy

Collaboration research gives us better primitives than emotion labels: joint attention and common ground.

People coordinate by pointing at the same object, using the same names, confirming what changed, and repairing misunderstandings. A review of multimodal collaboration research describes grounding as the process by which a group establishes and maintains a shared definition of the situation.

Read the collaboration review.

For Rover, that becomes a control policy:

  1. Identify the account, artifact, and state the group is discussing.
  2. Choose the least-powerful useful next step.
  3. Ask when uncertainty would change the action or the required authority.
  4. Act only in the active, authorized surface.
  5. Verify the resulting state, log, or receipt.
  6. Escalate when risk, policy, ambiguity, or repeated failure exceeds the boundary.

If a customer shares their screen, Rover can point or highlight without pretending it controls the device. If the shared product is Rover-enabled and the customer authorizes it, Rover can issue an account-scoped action through that product session and show the effect. If neither is true, Rover guides the human driver. Remote control is a permission state, not a meeting trick.

The name comes after the principle

The first version should not be an avatar with a better smile. It should be a participant that arrives prepared, sees the same authorized work, speaks when useful, can demonstrate or act when asked, and leaves behind an updated plan with receipts.

It should be able to say “I do not know why they looked away. I do know the integration failed, the group is looking at the wrong environment, and I need approval before changing it.”

That is the research direction we are pursuing. We call this direction Rover Live.

It is not a generally available feature announcement. We are looking for design partners who want to test one repeated customer workflow—onboarding, a technical validation, a demo, a recurring debug path, or a value review—with explicit permissions and measurable outcomes.

This work grows from the same gap we found in our study of 150 customer-engineering jobs: companies do not need another meeting transcript. They need continuity between what was said, what was done, what changed, and what happens next.

If your team has that workflow, show it to us. Put the repeated workflow in the “Anything specific you want to see?” field.

The complete meeting-context evidence review is public, including the claim boundaries behind this article.

More from the blog

On this page

  • A face cannot establish intent
  • What multimodal research succeeds at
  • The real meeting is multi-person, multi-device, and task-shaped
  • The legal and consent line
  • A better context model
  • Repair the work, not the person
  • Joint attention before autonomy
  • The name comes after the principle