OpenHuman - Architecture Overview
System Layers
Core Modules
1. Loader
Responsibility: Parse .ohb bundle files, upload assets to the GPU.
- Reads the binary
.ohbarchive (header + chunk list) - Decodes KTX2 textures → uploads as
WebGLTexture - Parses glTF mesh data → uploads as
WebGLBuffer(VBO + IBO) - Extracts skeleton, morph targets, and animation clips
- Emits
character:loadedevent when GPU upload is complete
Key classes: BundleParser, TextureUploader, GeometryUploader
2. Animator
Responsibility: Manage animation state and produce per-frame pose data.
- Implements a state machine with idle/talk/gesture states
- Supports blend trees for smooth transition between clips
- Evaluates 52 FACS morph target weights per frame
- Outputs a
PoseFrame: joint transforms + blendshape weights - Receives external poses from
StreamingClientand merges with local state
Key classes: AnimationGraph, StateMachine, BlendTree, MorphController
3. StreamingClient
Responsibility: Receive real-time animation data over the network.
- Opens a WebSocket (or HTTP chunked) connection to an animation server
- Decodes binary frames: 16-bit quantized joint data →
Float32Array - Maintains a jitter buffer to smooth latency spikes
- Pushes decoded
PoseFrameobjects into the Animator's input queue
Key classes: StreamingClient, JitterBuffer, FrameDecoder
4. Render Engine
Responsibility: Take a PoseFrame + scene state → produce final pixels.
Four sub-systems execute in order each frame:
| Sub-system | Role |
|---|---|
| Geometry Pipeline | GPU skinning (compute shader), frustum culling |
| Shadow Map | Render depth from light POV, PCF filtering |
| Material System | PBR shading, SSS pass, draw calls |
| Post-Process Stack | Bloom → DoF → ACES tonemapping → FXAA |
5. WebGL 2.0 Context Manager
Responsibility: Own and manage the raw WebGL context.
- Created once at
new OpenHuman({ canvas })initialization - Manages WebGL extension detection and capability flags
- Handles context loss/restore events
- Provides a thin abstraction layer (
GpuDevice) used by all sub-systems - rawWebGLRenderingContextis never exposed in the public API
Data Flow: Asset Loading
Data Flow: Per-Frame Render Loop
Threading Model
OpenHuman runs on a single main thread by default, with optional worker offloading:
| Task | Thread | Notes |
|---|---|---|
| Render loop | Main thread | requestAnimationFrame |
| Asset parsing | Web Worker | Offloaded via BundleParser worker |
| WebSocket I/O | Main thread | Browser handles I/O async |
| Frame decoding | Web Worker | FrameDecoder runs in worker, posts PoseFrame |
| GPU commands | Main thread | WebGL requires main thread (no OffscreenCanvas by default) |
OffscreenCanvas support (Chrome only): pass
offscreen: truetonew OpenHuman()to move the render loop to a dedicated worker thread. See GPU Optimization Guide for details.
Public API Surface
The SDK exposes a minimal API surface. All internal sub-systems are private.
class OpenHuman {
// Lifecycle
constructor(config: OpenHumanConfig)
loadCharacter(url: string): Promise<void>
destroy(): void
// Playback
play(animation: string, options?: PlayOptions): void
stop(): void
applyPose(pose: PoseFrame): void
// Morphs
setMorphWeight(name: string, weight: number): void
setMorphWeights(weights: Record<string, number>): void
// Configuration
setQuality(quality: "high" | "medium" | "low"): void
setFPS(fps: number): void
// Events
on(event: string, handler: Function): void
off(event: string, handler: Function): void
// Debug
getStats(): RenderStats
}Key Design Decisions
Why pure WebGL 2.0 (no Three.js)? Three.js and Babylon.js are general-purpose engines with significant overhead (scene graph, physics, audio, etc.) that OpenHuman doesn't need. A purpose-built renderer for digital humans allows tighter control over the render pipeline, SSS implementation, and GPU memory layout - resulting in a ≤200KB bundle vs. 500KB+ for a general engine.
Why .ohb instead of raw glTF?
The .ohb format pre-processes and pre-optimizes assets for the OpenHuman pipeline: KTX2 textures are already in GPU-native compressed formats, morph target deltas are pre-computed, and the skeleton is already in OpenHuman's joint order. This eliminates runtime parsing overhead and enables faster load times.
Why 16-bit quantization for streaming? Full 32-bit floats for all joints would require ~2KB per frame at 60fps = ~120KB/s per character. 16-bit quantization halves this to ~60KB/s with imperceptible quality loss for animation data within human joint range-of-motion limits.
Next Steps
- Render Pipeline Deep Dive - detailed pass-by-pass breakdown
.ohbBundle Format Spec - binary layout and chunk types- Animation Graph Reference - state machine configuration
- Streaming Protocol Spec - WebSocket frame format
- GPU Optimization Guide - memory budgets and profiling