A single photo is now enough to hold a conversation with an animated version of it. Runway Characters takes one reference image, a headshot, a cartoon, a fantasy creature, or a brand mascot, and turns it into a video agent that listens, looks, and talks back in close to real time, with no fine-tuning and no per-character training step. Runway described the system on May 4, 2026, and the pitch is that the hardest part of a talking avatar, getting one built at all, collapses to uploading an image.
The performance claims are specific. Runway says the agent runs at 24 frames per second in HD and spends 37 milliseconds of model time per frame, inside a 42-millisecond budget. From the moment a person stops speaking to the character's first response, the company measures a 1.75-second turnaround, which it breaks down as roughly 1,185 milliseconds for the voice component, 567 milliseconds for the video pipeline, and about 400 milliseconds of network round-trip. Those numbers are Runway's own, measured on its infrastructure, so they read as best-case rather than guaranteed, but they are precise enough to hold the company to.
What separates this from a lip-sync clip is that the character is wired to do things during a session. Runway says the agent can see through a webcam or a shared screen, use a custom voice built from a prompt or cloned from a short sample, and call tools, meaning it can trigger a UI action or hit a backend function such as a customer's order-status endpoint. Teams can attach a knowledge base of text or Markdown for domain answers, embed the character in a web app with a single line of code, or drop it into a Zoom, Google Meet, or Teams call. The company frames the uses as tutoring, product demos, games, design feedback, and customer calls.
Getting a diffusion model to generate video fast enough to argue with is the engineering story here. Runway Characters is built on GWM-1, the company's general world model, and generates frames autoregressively one at a time with a streaming pipeline rather than rendering a finished clip and playing it back. To keep latency inside the frame budget, Runway credits device sharding with pipeline and tensor-parallel diffusion, aggressive KV-cache eviction and compression, CUDA Graphs to cut kernel-launch overhead, and fused Triton kernels for attention and matrix math. The agent is available through the Runway API and the company's web and mobile apps.
Real-time avatars have been the missing piece in generative video. Text-to-video systems, Runway's own Gen-series among them, produce polished clips but not something that reacts to you mid-sentence, and that gap is what keeps most AI characters feeling like recordings. Whether an image-only pipeline holds up outside a launch demo is the open question, since a live agent has to survive interruptions, awkward lighting, and off-script prompts that a curated reel never shows. If it does, the low bar to entry, one picture, moves the interesting problem from making a character look right to deciding what a character built from a single stranger's photo should be allowed to say and do.













