Real-time affective co-speech motion for humanoid robots

SocialHumanoid

Towards Expressive Humanoid Behavior via One-Step Co-Speech Motion Generation

Given streaming speech and an affective condition, SocialHumanoid synthesizes emotionally expressive full-body motion, retargets it online, and executes it on a physical humanoid robot.

View demos Read results Paper (coming soon) arXiv (coming soon) Code (coming soon) Dataset (coming soon)

Overview

Speech-aligned motion with visible affect and robot-ready execution.

Humanoid co-speech motion must do more than look plausible on a virtual skeleton. It needs to stay synchronized with speech, remain executable under robot constraints, and preserve distinguishable emotional styles after retargeting and control.

Expressive humanoid behavior generated from spoken response and affective condition.
SocialHumanoid converts a spoken response and an affective condition into continuous full-body co-speech motion, then retargets and executes it online on a physical humanoid.
3.929 FGD on BEAT2, best among compared methods
0.0281s AIST for a 128-frame window on RTX 3090
1 NFE One-step improved MeanFlow inference
4h / 8 AffectMoCap hours and emotion labels

System Pipeline

From dialogue and speech to controllable expressive humanoid behavior.

The runtime system connects response generation, speech synthesis, affect-conditioned motion generation, online retargeting, and whole-body control.

SocialHumanoid pipeline for affective supervision, one-step co-speech motion generation, humanoid deployment, and interactive communication.
Streaming speech and affective conditions drive an improved MeanFlow generator in SMPL-X space. The generated motion is retargeted into humanoid joint targets and tracked by a whole-body controller for synchronized robot execution.

Robot Demos

Continuous co-speech motion in real-world robot scenes.

Demonstrations cover BEAT2 benchmark comparisons, interactive dialogue, emotion-conditioned motion, multilingual speech, and long-horizon execution.

BEAT2 Dataset Comparison

Motion-space and robot-space comparisons across representative baselines.

Each case compares the same BEAT2 speech clip in the digital human motion space and after retargeting to the physical humanoid robot.

Case I

Sample 2_scott_0_1_1

Case II

Sample 2_scott_0_2_2

AffectMoCap Dataset

Professional actor motion capture for affective body expression.

AffectMoCap contains synchronized speech, full-body SMPL-X motion, transcripts, speaker metadata, and emotion annotations across eight affective labels.

What it captures

  • Two professional actors, one male and one female.
  • Eight affective labels: neutral, happiness, anger, sadness, surprise, fear, disgust, and enthusiasm.
  • Full-body marker motion, hand motion-capture gloves, and synchronized 48 kHz speech audio.
  • Processed 16 kHz audio and 333D SMPL-X motion frames compatible with BEAT2-style training.
Representative AffectMoCap skeleton and SMPL-X mesh samples from two actors.
Representative AffectMoCap samples from two professional actors.
Duration distribution of AffectMoCap by actor gender and emotion label.
Duration distribution by actor gender and emotion category.
OptiTrack camera, marker suit, hand motion-capture gloves, and microphone kit used for AffectMoCap.
Capture setup with optical tracking, gloves, and wireless audio.

Method

A co-speech generator wrapped in deployment-aware robot execution.

01

Part-wise latent motion

Upper body, hands, lower body, and translation are encoded with residual vector quantized autoencoders, preserving detail while keeping online inference compact.

02

Affect-conditioned generation

Low-latency acoustic features are fused with emotion and speaker embeddings, then denoised by cross-attention and spatial-temporal Transformer blocks.

03

Improved MeanFlow

One-step interval-velocity sampling synthesizes each motion window fast enough for interactive co-speech deployment.

04

Retargeting and control

Generated SMPL-X motion is converted to humanoid references with online retargeting, smoothed, queued, and tracked by a whole-body controller.

Improved MeanFlow module for one-step latent motion generation.
Improved MeanFlow predicts interval velocity in latent motion space for one-step real-time motion synthesis.

Results

Better motion realism with competitive synchrony and fast inference.

SocialHumanoid is evaluated on BEAT2, AffectMoCap, perceptual studies, ablations, and physical humanoid deployment.

BEAT2 Quantitative Evaluation

Method FGD (down) BC (up) Diversity (up) NFE (down)
EMAGE5.5120.77213.062
MambaTalk5.3660.78113.052
SynTalker4.6870.73612.431000
GestureLSM4.0880.71413.248
LiveGesture4.5700.79413.9132
SocialHumanoid3.9290.77012.361

Runtime and Robot Modules

AIST on RTX 3090 0.0281s
Motion generation on RTX 4090 0.006s
Online retargeting 1.248s
Controller execution 0.008s

Emotion and Style Recognition

AffectMoCap, audio + motion
92.4%
AffectMoCap, motion only
88.6%
SocialHumanoid, audio + generated motion
86.3%
SocialHumanoid, generated motion only
82.7%

Ablation Summary

Variant FGD BC NFE
w/o temporal attention4.6480.7871
w/o spatial attention5.0310.7831
w/o affective condition4.7300.7681
MeanFlow retraining4.0780.7691
Shortcut retraining4.2410.7288
Full model3.9290.7701