OpenHuman - Animation Graph Reference
Overview
The Animation Graph is a runtime state machine that controls which animation clips play, how they blend together, and when transitions occur. It is the core of the character's local animation logic - responsible for everything from idle breathing to gesture playback - and acts as a merge point for externally streamed poses.
flowchart LR
A[Local clips] --> B
C[Stream poses] --> B
D[Morph API] --> E
subgraph B[Animation Graph]
direction TB
F[State Machine]
G[Blend Tree]
H[Stream Merger]
E[Morph Controller]
F --> G
end
G --> I[PoseFrame]
I --> J[GPU Skinning]
H --> G
E --> K[MorphWeights]
K --> L[GPU Morphing]Concepts
| Concept | Description |
|---|---|
| State | A named playback context - plays one clip or a blend of clips |
| Transition | Moves the graph from one state to another, with optional blend-out/blend-in |
| Parameter | A named value (float, bool, trigger) used to drive transitions |
| Blend Tree | A node graph within a state that blends multiple clips by parameter values |
| Layer | A stack of state machines that blend additively (e.g. body + face layers) |
| Stream Merger | Combines the local graph output with incoming streamed joint data |
Defining an Animation Graph
Graphs are defined in JSON and loaded alongside the .ohb bundle, or passed directly to the SDK:
await human.loadCharacter("./character.ohb", {
animationGraph: "./character.graph.json",
})
// - or inline -
human.setAnimationGraph(graphDefinition)Minimal Graph Example
{
"version": "1.0",
"entry_state": "idle",
"parameters": {
"isTalking": { "type": "bool", "default": false },
"talkSpeed": { "type": "float", "default": 1.0 },
"doNod": { "type": "trigger" }
},
"states": {
"idle": {
"clip": "idle",
"loop": true,
"transitions": [
{ "to": "talking", "condition": "isTalking == true", "duration_ms": 200 },
{ "to": "nod", "condition": "doNod", "duration_ms": 100 }
]
},
"talking": {
"clip": "talk",
"loop": true,
"speed_param": "talkSpeed",
"transitions": [{ "to": "idle", "condition": "isTalking == false", "duration_ms": 300 }]
},
"nod": {
"clip": "nod",
"loop": false,
"transitions": [{ "to": "idle", "condition": "clip_finished", "duration_ms": 150 }]
}
}
}Parameters
Parameters are typed values that drive state transitions and blend tree weights.
Parameter Types
| Type | Values | Use Case |
|---|---|---|
bool | true / false | State toggles (e.g. isTalking) |
float | any f32 | Blend weights, speed multipliers |
trigger | consumed on use | One-shot events (e.g. doNod, doBlink) |
int | integer | Discrete state selection |
Setting Parameters at Runtime
// bool
human.setParam("isTalking", true)
// float
human.setParam("talkSpeed", 1.25)
// trigger - fires once and auto-resets
human.triggerParam("doNod")
human.triggerParam("doBlink")
// int
human.setParam("emotionIndex", 2)Reading Parameters
const talking = human.getParam("isTalking") // boolean
const speed = human.getParam("talkSpeed") // numberStates
Each state is a named node in the graph that plays one or more animation clips.
State Properties
interface StateDefinition {
clip?: string // single clip name from .ohb bundle
blend_tree?: BlendTree // mutually exclusive with `clip`
loop?: boolean // default: true
speed?: number // playback speed multiplier (default: 1.0)
speed_param?: string // float param to drive speed (overrides `speed`)
offset?: number // start time offset in seconds
transitions: Transition[]
}Entry State
The entry_state field in the graph definition specifies which state the graph enters on load. The character begins playing this state immediately after loadCharacter() resolves.
Clip Finished Condition
For non-looping states, the special condition "clip_finished" fires when the clip reaches its last frame. Use this to return to idle after gestures:
{
"to": "idle",
"condition": "clip_finished",
"duration_ms": 150
}Transitions
Transitions define when and how the graph moves between states.
Transition Properties
interface Transition {
to: string // target state name
condition: string // condition expression (see below)
duration_ms?: number // crossfade duration (default: 0 = instant)
blend_curve?: BlendCurve // easing for crossfade (default: "ease_in_out")
interrupt?: boolean // can this transition interrupt an active transition?
priority?: number // higher = evaluated first (default: 0)
}Condition Expressions
Conditions are evaluated every frame. Supported syntax:
// bool parameter
"isTalking == true"
"isTalking == false"
"isTalking" // shorthand for == true
// float comparisons
"talkSpeed > 1.5"
"emotionBlend <= 0.2"
"emotionBlend >= 0.5 && isTalking == false"
// trigger (consumed on evaluation)
"doNod"
"doBlink"
// clip state
"clip_finished"
"clip_time > 0.8" // proportional clip time [0.0โ1.0]
// logical operators
"isTalking && talkSpeed > 1.0"
"!isTalking || emotionBlend < 0.1"
Blend Curves
| Name | Description |
|---|---|
linear | Constant crossfade rate |
ease_in | Slow start, fast finish |
ease_out | Fast start, slow finish |
ease_in_out | Slow start and finish (default) |
hold_then_cut | Hold source until midpoint, then cut |
Transition Interrupts
By default (interrupt: false), a transition that is already in progress cannot be interrupted by a new transition. Set interrupt: true on high-priority transitions (e.g. blink) to allow them to override ongoing crossfades:
{
"to": "blink",
"condition": "doBlink",
"duration_ms": 80,
"interrupt": true,
"priority": 10
}Blend Trees
A Blend Tree is used inside a state to blend multiple clips together based on parameter values. This allows smooth, data-driven transitions between animation variations within a single state.
1D Blend Tree
Blends clips linearly along a single float parameter axis:
"talking": {
"blend_tree": {
"type": "1d",
"param": "talkSpeed",
"nodes": [
{ "clip": "talk_slow", "value": 0.5 },
{ "clip": "talk_normal", "value": 1.0 },
{ "clip": "talk_fast", "value": 1.8 }
]
}
}At talkSpeed = 1.3, the graph blends talk_normal (70%) with talk_fast (30%).
2D Blend Tree
Blends clips across a 2D float parameter space. Useful for body lean direction:
"idle": {
"blend_tree": {
"type": "2d",
"param_x": "leanX",
"param_y": "leanY",
"nodes": [
{ "clip": "idle_center", "x": 0.0, "y": 0.0 },
{ "clip": "idle_lean_left", "x": -1.0, "y": 0.0 },
{ "clip": "idle_lean_right", "x": 1.0, "y": 0.0 },
{ "clip": "idle_lean_forward","x": 0.0, "y": 1.0 },
{ "clip": "idle_lean_back", "x": 0.0, "y": -1.0 }
]
}
}2D blending uses barycentric interpolation over the Delaunay triangulation of the node positions.
Direct Blend Tree
Blends N clips with explicit weight parameters - useful when each weight is independently controlled (e.g. emotion blending):
"emotional": {
"blend_tree": {
"type": "direct",
"nodes": [
{ "clip": "emotion_neutral", "weight_param": "weightNeutral" },
{ "clip": "emotion_happy", "weight_param": "weightHappy" },
{ "clip": "emotion_sad", "weight_param": "weightSad" },
{ "clip": "emotion_angry", "weight_param": "weightAngry" }
]
}
}Weights are automatically normalized so they sum to 1.0. Set all weights programmatically:
human.setParam("weightNeutral", 0.0)
human.setParam("weightHappy", 0.8)
human.setParam("weightSad", 0.2)
human.setParam("weightAngry", 0.0)Layers
The animation graph supports multiple layers that blend additively. Each layer has its own state machine, and layer outputs are combined top-to-bottom.
Use Cases
| Layer | Contents | Blend Mode |
|---|---|---|
base | Full-body locomotion / idle | Override (base) |
upper_body | Upper body gestures | Override (masked) |
face | Facial expressions | Additive |
breathing | Subtle chest movement | Additive |
Layer Configuration
{
"layers": [
{
"name": "base",
"weight": 1.0,
"blend_mode": "override",
"mask": null,
"entry_state": "idle",
"states": { ... }
},
{
"name": "face",
"weight": 1.0,
"blend_mode": "additive",
"mask": ["head", "jaw", "eye_l", "eye_r"],
"entry_state": "face_neutral",
"states": { ... }
}
]
}Layer Masks
A joint mask restricts a layer to only affect specific joints. Other joints retain the pose from layers below them. The mask is an array of joint names - all children of named joints are automatically included.
Setting "mask": ["head"] includes head + all 20+ facial joints automatically.
Layer Weight
human.setLayerWeight("face", 0.5) // blend face layer at 50% influence
human.setLayerWeight("base", 1.0) // base layer always full weightMorph Controller
The Morph Controller manages the 52 FACS blendshape weights independently of the joint animation system.
Direct API
// Set a single weight (0.0 โ 1.0)
human.setMorphWeight("mouthSmileLeft", 0.8)
human.setMorphWeight("mouthSmileRight", 0.8)
human.setMorphWeight("cheekSquintLeft", 0.4)
// Set many weights at once (efficient - single GPU upload)
human.setMorphWeights({
jawOpen: 0.3,
mouthFunnel: 0.1,
browInnerUp: 0.2,
})
// Reset all morphs to 0
human.resetMorphWeights()Morph Animations
Morph weights can be animated over time using the built-in tweening API:
// Animate to target weights over 300ms
human.animateMorphWeights({ mouthSmileLeft: 0.9, mouthSmileRight: 0.9, cheekSquintLeft: 0.5 }, { duration: 300, easing: "ease_out" })Procedural Blink
The SDK includes a built-in procedural blink controller that fires automatically at randomized human-realistic intervals:
human.setBlink({
enabled: true,
intervalMin: 2500, // ms between blinks (minimum)
intervalMax: 6000, // ms between blinks (maximum)
duration: 180, // ms for a full blink cycle (down + up)
})The blink controller drives eyeBlinkLeft and eyeBlinkRight morph weights directly, with a slight random timing offset between left and right eye for natural asymmetry.
Viseme / Lip Sync API
For TTS-driven lip sync, morph weights are typically driven via the streaming protocol (MORPH frames). However, the SDK also exposes a client-side viseme API for audio-driven sync:
// Apply a viseme blend over a given time window
human.applyViseme({
viseme: "PP", // viseme name (see Viseme Table below)
weight: 0.85, // peak weight
onset_ms: 0, // time from now when weight starts rising
peak_ms: 80, // time from now at peak weight
offset_ms: 180, // time from now when weight returns to 0
})Viseme to FACS Mapping Table
| Viseme | Phonemes | Primary Morphs | Secondary Morphs |
|---|---|---|---|
PP | p, b, m | mouthClose=1.0 | mouthPressLeft=0.3, mouthPressRight=0.3 |
FF | f, v | mouthLowerDownLeft=0.4, mouthLowerDownRight=0.4 | mouthUpperUpLeft=0.2 |
TH | th | tongueOut=0.6 | jawOpen=0.1 |
DD | t, d | tongueOut=0.3 | jawOpen=0.2 |
KK | k, g | jawOpen=0.4 | mouthFunnel=0.1 |
CH | ch, sh | mouthFunnel=0.6 | jawOpen=0.3 |
SS | s, z | mouthShrugUpper=0.2 | jawOpen=0.15 |
NN | n, l | jawOpen=0.25 | tongueOut=0.1 |
RR | r | mouthFunnel=0.4 | mouthPucker=0.2 |
AA | a, ah | jawOpen=0.7 | mouthSmileLeft=0.1, mouthSmileRight=0.1 |
EE | e, ee | mouthSmileLeft=0.5, mouthSmileRight=0.5 | jawOpen=0.3 |
II | ih | mouthSmileLeft=0.3, mouthSmileRight=0.3 | jawOpen=0.2 |
OO | o, ow | mouthFunnel=0.8 | mouthPucker=0.3, jawOpen=0.4 |
UU | u, oo | mouthPucker=0.9 | mouthFunnel=0.3 |
sil | silence | (all at 0) | - |
Stream Merger
When a StreamingClient is connected, incoming PoseFrame data from the server is merged with the local animation graph output via the Stream Merger.
Merge Modes
human.setStreamMergeMode("override") // streaming pose replaces local graph entirely
human.setStreamMergeMode("additive") // streaming pose is added on top of local graph
human.setStreamMergeMode("masked") // streaming overrides specific joints onlyMasked Merge
The most common mode for TTS-driven talking heads - the server streams facial joint poses while the local graph drives body animation:
human.setStreamMergeMode('masked', {
streamMask: ['head', 'jaw', 'eye_l', 'eye_r', 'neck'], // joints from stream
localMask: ['hips', 'spine_01', 'spine_02', 'chest', // joints from local graph
'shoulder_l', 'upper_arm_l', ...]
})Stream Fallback
When the stream disconnects, the merger automatically fades back to the local animation graph over streamFallbackDuration ms:
human.connectStream(url, {
streamFallbackDuration: 500, // ms to blend back to local graph on disconnect
streamFallbackState: "idle", // which state to transition to on disconnect
})Runtime Control
Playback Control
// Manually play a state (bypasses transition conditions)
human.play("talking")
human.play("idle")
human.play("nod")
// Stop playback and hold last pose
human.pause()
// Resume from paused
human.resume()
// Seek within current clip (0.0 โ 1.0)
human.seekClip(0.5)Graph Inspection (Debug)
const state = human.getGraphState()
// {
// activeState: 'talking',
// activeLayer: 'base',
// transitionProgress: 0.0, // 0.0 = not transitioning
// clipTime: 0.73, // 0.0โ1.0 position in current clip
// blendWeights: { talk_slow: 0.0, talk_normal: 0.7, talk_fast: 0.3 },
// morphWeights: { jawOpen: 0.3, mouthFunnel: 0.1, ... }
// }Full Graph Schema Reference
interface AnimationGraphDefinition {
version: "1.0"
entry_state: string
parameters: Record<string, ParameterDef>
states?: Record<string, StateDef> // single-layer shorthand
layers?: LayerDef[] // multi-layer
}
interface ParameterDef {
type: "bool" | "float" | "int" | "trigger"
default?: boolean | number
min?: number // float/int only
max?: number // float/int only
}
interface StateDef {
clip?: string
blend_tree?: BlendTreeDef
loop?: boolean
speed?: number
speed_param?: string
offset?: number
transitions: TransitionDef[]
}
interface TransitionDef {
to: string
condition: string
duration_ms?: number
blend_curve?: "linear" | "ease_in" | "ease_out" | "ease_in_out" | "hold_then_cut"
interrupt?: boolean
priority?: number
}
interface BlendTreeDef {
type: "1d" | "2d" | "direct"
param?: string // 1d
param_x?: string // 2d
param_y?: string // 2d
nodes: BlendNodeDef[]
}
interface BlendNodeDef {
clip: string
value?: number // 1d position
x?: number // 2d position
y?: number // 2d position
weight_param?: string // direct blend
}
interface LayerDef {
name: string
weight: number
blend_mode: "override" | "additive"
mask?: string[]
entry_state: string
states: Record<string, StateDef>
}Best Practices
Keep the base layer simple. Use 2โ4 states maximum in the base layer (idle, talk, gesture, expression). Complex logic belongs in blend trees, not as many states.
Use triggers for one-shots. Gestures like nods, blinks, and laughs should use trigger parameters + clip_finished return transitions - never booleans.
Layer facial animation separately. Always put facial joint animation on a dedicated face layer with a joint mask. This allows body and face animation to evolve independently.
Stream only what changes. In masked merge mode, only stream the joints that the server is actively controlling. Streaming unused joints wastes bandwidth and can introduce subtle pose artifacts.
Pre-warm transitions. For latency-sensitive applications, call human.preloadState('talking') on page load so the transition is instant when isTalking fires.
Next Steps
- WASM Runtime Docs - animation graph evaluation internals
- Streaming Protocol Spec - how POSE frames feed into the Stream Merger
- TTS Integration Guide - driving morph weights from audio
- GPU Optimization Guide - morph target GPU budget