OpenHuman - Architecture Overview


System Layers


Core Modules

1. Loader

Responsibility: Parse .ohb bundle files, upload assets to the GPU.

  • Reads the binary .ohb archive (header + chunk list)
  • Decodes KTX2 textures → uploads as WebGLTexture
  • Parses glTF mesh data → uploads as WebGLBuffer (VBO + IBO)
  • Extracts skeleton, morph targets, and animation clips
  • Emits character:loaded event when GPU upload is complete

Key classes: BundleParser, TextureUploader, GeometryUploader


2. Animator

Responsibility: Manage animation state and produce per-frame pose data.

  • Implements a state machine with idle/talk/gesture states
  • Supports blend trees for smooth transition between clips
  • Evaluates 52 FACS morph target weights per frame
  • Outputs a PoseFrame: joint transforms + blendshape weights
  • Receives external poses from StreamingClient and merges with local state

Key classes: AnimationGraph, StateMachine, BlendTree, MorphController


3. StreamingClient

Responsibility: Receive real-time animation data over the network.

  • Opens a WebSocket (or HTTP chunked) connection to an animation server
  • Decodes binary frames: 16-bit quantized joint data → Float32Array
  • Maintains a jitter buffer to smooth latency spikes
  • Pushes decoded PoseFrame objects into the Animator's input queue

Key classes: StreamingClient, JitterBuffer, FrameDecoder


4. Render Engine

Responsibility: Take a PoseFrame + scene state → produce final pixels.

Four sub-systems execute in order each frame:

Sub-systemRole
Geometry PipelineGPU skinning (compute shader), frustum culling
Shadow MapRender depth from light POV, PCF filtering
Material SystemPBR shading, SSS pass, draw calls
Post-Process StackBloom → DoF → ACES tonemapping → FXAA

5. WebGL 2.0 Context Manager

Responsibility: Own and manage the raw WebGL context.

  • Created once at new OpenHuman({ canvas }) initialization
  • Manages WebGL extension detection and capability flags
  • Handles context loss/restore events
  • Provides a thin abstraction layer (GpuDevice) used by all sub-systems - raw WebGLRenderingContext is never exposed in the public API

Data Flow: Asset Loading


Data Flow: Per-Frame Render Loop


Threading Model

OpenHuman runs on a single main thread by default, with optional worker offloading:

TaskThreadNotes
Render loopMain threadrequestAnimationFrame
Asset parsingWeb WorkerOffloaded via BundleParser worker
WebSocket I/OMain threadBrowser handles I/O async
Frame decodingWeb WorkerFrameDecoder runs in worker, posts PoseFrame
GPU commandsMain threadWebGL requires main thread (no OffscreenCanvas by default)

OffscreenCanvas support (Chrome only): pass offscreen: true to new OpenHuman() to move the render loop to a dedicated worker thread. See GPU Optimization Guide for details.


Public API Surface

The SDK exposes a minimal API surface. All internal sub-systems are private.

class OpenHuman {
    // Lifecycle
    constructor(config: OpenHumanConfig)
    loadCharacter(url: string): Promise<void>
    destroy(): void
 
    // Playback
    play(animation: string, options?: PlayOptions): void
    stop(): void
    applyPose(pose: PoseFrame): void
 
    // Morphs
    setMorphWeight(name: string, weight: number): void
    setMorphWeights(weights: Record<string, number>): void
 
    // Configuration
    setQuality(quality: "high" | "medium" | "low"): void
    setFPS(fps: number): void
 
    // Events
    on(event: string, handler: Function): void
    off(event: string, handler: Function): void
 
    // Debug
    getStats(): RenderStats
}

Key Design Decisions

Why pure WebGL 2.0 (no Three.js)? Three.js and Babylon.js are general-purpose engines with significant overhead (scene graph, physics, audio, etc.) that OpenHuman doesn't need. A purpose-built renderer for digital humans allows tighter control over the render pipeline, SSS implementation, and GPU memory layout - resulting in a ≤200KB bundle vs. 500KB+ for a general engine.

Why .ohb instead of raw glTF? The .ohb format pre-processes and pre-optimizes assets for the OpenHuman pipeline: KTX2 textures are already in GPU-native compressed formats, morph target deltas are pre-computed, and the skeleton is already in OpenHuman's joint order. This eliminates runtime parsing overhead and enables faster load times.

Why 16-bit quantization for streaming? Full 32-bit floats for all joints would require ~2KB per frame at 60fps = ~120KB/s per character. 16-bit quantization halves this to ~60KB/s with imperceptible quality loss for animation data within human joint range-of-motion limits.


Next Steps