A whole new look.
Change the character.
Use a prompt to explore character styles while your camera supplies the pose and expression.
Turn a live camera into a creative canvas. Transform your character with AI that runs on your own machine.
01 / THE POSSIBILITIES
Explore a new look. Build a character. Give your next creative experiment a world of its own.
Use a prompt to explore character styles while your camera supplies the pose and expression.
Run core AI inference on your own GPU. Open the studio in your browser and keep your workflow close.
Start with your camera, edit existing footage, or queue multiple tasks in a single session.
THE TRANSFORMATION / IN MOTION
Watch appearance edits follow the source motion, in a presentation redesigned for Dash.
02 / YOUR CREATIVE WORKFLOW
From a blank prompt to a character experiment, all in one local workspace.
Explore the studioSet up the models on your GPU machine, open the browser interface, and prepare your preferred configuration.
LOCAL WORKSPACEWrite your character direction. Try a different material, an animated look, or a softer creative style.
PROMPT-BASED EDITINGEnable your camera, frame the shot, and start editing. Reset, change the prompt, and explore again.
CREATE. RESET. REPEAT.03 / UNDER THE HOOD
Appearance, motion, and temporal context come together in a streaming video diffusion pipeline.
Wan2.2-Animate
Explicit pose and expression signals
Chunk-by-chunk streaming cache
Only the reference appearance is edited synthetically. A reversed prompt then trains reconstruction of the original video. Clean-reference conditioning, body pose, and facial features separate appearance from motion.
Each target chunk sees its own chunk, the clean reference, and preceding clean context. Future context and other noisy chunks are masked. Flow-matching loss is applied only to targets.
Training completes the denoising rollout and backpropagates at the sampled step. A retained attention sink and rolling local cache support continued generation; fixed RoPE and FPSA preserve the intended temporal context.
The figures and animation explain the published training design. The studio runtime uses pretrained adapters; training code is not currently included. Source paper: Figures 3 and 5, Appendix B.1 ↗
Three stages. One guided walkthrough.
1:26 · 1080p · Silent animation · Use fullscreen for the details.
The research, explained by an AI presenter.
2:01 · 1080p · English captions available with CC · Fullscreen for the details.
Welcome to Dash. Let’s unpack the research pipeline behind its local video-editing workflow. This diagram describes three training stages, rather than a live recording of a model running.
First, appearance and motion are separated. Only the reference image is edited synthetically, and a reverse instruction trains reconstruction of the original video. Body-pose signals and facial features guide movement, while the reference and editing instruction jointly control appearance.
Next, the model adapts to streaming. Each chunk can attend throughout itself, to the clean reference, and to preceding clean context. Future context and other noisy chunks are masked. Training applies the flow-matching loss to target chunks only.
The third stage is aligned self-rollout distillation. The model completes both denoising steps, while gradients pass through the sampled step. An extra clean cache pass builds the retained attention sink from the first generated chunk. Subsequent chunks pass their final denoising-step keys and values to following chunks. Frozen real-score and trainable fake-score diffusion models provide the distillation signal.
A retained attention sink and rolling local cache manage context. Fixed rotary positions and first-frame-preserved sparse attention help maintain appearance and temporal consistency. The numbers here identify latent context positions, not ordinary video-frame numbers.
Dash brings pretrained components from this pipeline into a local studio workflow. Actual output quality and speed depend on the input, settings, and hardware.
Loading diagram…Open diagram separately ↗
Figure 01. Reference latents, motion adapters, causal attention, and the rolling cache connect the three training stages.
Loading diagram…Open diagram separately ↗
Figure 02. Full-rollout alignment keeps training context consistent with inference. The aligned updates use two network evaluations; the comparison also shows cache-building forwards.
THE LOCAL RUNTIME
Choose the balance of image size, memory use, and throughput that fits your GPU.
FILE INFERENCE / H100 REFERENCE
For file inference, move weights or the rolling KV cache to CPU memory to reduce GPU memory use, with a speed tradeoff.
FILE INFERENCE · H100 · FP8 · FAST DECODE · NO COMPILE| Memory mode | 672 × 384 | 832 × 480 |
|---|---|---|
| GPU resident | 31.4 GiBBaseline | 38.4 GiBBaseline |
| CPU weights | 17.5 GiB1.02× slower | 24.5 GiB1.02× slower |
| CPU weights + KV | 14.4 GiB1.71× slower | 15.0 GiB1.57× slower |
These are reported source measurements, rather than Dash-specific validation. Capture FPS is a camera upload target, not guaranteed AI output speed. Low-memory modes and torch.compile are mutually exclusive.
SYSTEM REQUIREMENTS
Live camera and file editing have different memory paths. Choose hardware for the mode you intend to use.
The current camera interface uses compilation and GPU-resident weights and cache. CPU offload is not exposed here. Allow extra headroom for pose, decoding, and compilation.
A live-camera minimum has not been published. 48 GB is a provisional workstation target, not a validated minimum.Start at 672 × 384 with FP8, Flash-VAED, CPU weights + KV offload, and compilation disabled. The lowest reported footprint for this configuration is 14.4 GiB on H100.
16 GB is inferred from that footprint. A 16 GB consumer GPU has not been validated here; 24 GB provides more headroom.| Component | Documented requirement | Practical planning target |
|---|---|---|
| GPU architecture | NVIDIA CUDA. FP8 requires SM 8.9 or newer. | Ada or Hopper for the pinned FP8 workflow. RTX 4090 is an Ada candidate; H100 is the reported benchmark GPU. |
| System RAM | No validated RAM minimum published. CPU offload needs additional host memory. | 64 GB RAM to plan model loading and offload; 128 GB for more build and multitasking headroom. Estimates require validation. |
| CPU | No validated core-count minimum published. | Modern 8-core x86-64 CPU or better. Build tools and preprocessing use the CPU; generation relies on the GPU. |
| Storage | Space for base weights, adapters, dependencies, caches, and outputs. | 150 GB free SSD space as a setup allowance. Reserve additional space for downloaded caches and generated videos; not a measured minimum. |
| Operating system | Linux is the primary build path for the pinned attention library. Native Windows builds need additional testing. | Ubuntu 22.04 LTS, 64-bit is a CUDA 12.4-compatible starting point, not a validated Dash installation. Windows/WSL2 and webcam forwarding need separate testing; macOS, Apple GPUs, AMD, and CPU-only use are not advertised as supported. |
| Python & PyTorch | Audited stack: Python 3.10, PyTorch 2.6.0, torchvision 0.21.0. | Use the pinned CUDA-enabled packages in an isolated environment. |
| CUDA & compiler | CUDA toolkit 12.3+ and GCC 10+ for FastVideo kernels. The dependency stack was audited with CUDA 12.4. | CUDA 12.4 with a compatible NVIDIA driver. Match the driver, toolkit, PyTorch build, and GPU architecture. |
| GPU driver | The NVIDIA driver must support the installed CUDA build. | For the CUDA 12.4 GA stack, use the corresponding Linux driver level 550.54.14 or newer. CUDA minor-version compatibility has separate rules; confirm the complete setup. |
| Kernel packages | FlashAttention 2.7.2.post1, FastVideo kernel 0.3.0, and Ninja. | Limit compilation jobs when RAM is constrained. Kernel builds are part of initial setup. |
| Browser & camera | Browser camera access on localhost or HTTPS. Camera permission is required for live mode. | A working webcam and current desktop browser. Node.js and npm are needed to build the interface; FFmpeg is needed for video preprocessing. |
| Internet | Needed initially to download models and install dependencies. | Once installed, core inference runs locally. Using a remote GPU requires a connection to that machine. |
VRAM capacity alone does not establish compatibility or live speed. The 14.4–15.0 GiB numbers describe file inference with FP8 and CPU offload, not the webcam interface. Ampere cards such as RTX 3090 do not meet this FP8 path’s SM 8.9 requirement. Newer architectures still need compatible builds of the pinned kernels.
Check NVIDIA compute capabilityAttention library platform requirementsCUDA 12.4 Linux environmentCUDA driver compatibilityThe VAE encodes frames into compact representations called latents and decodes generated latents into images.
Additional model weights that adapt the base model for editing, acceleration, and streaming.
Stored attention keys and values that carry context from previous video chunks into the current chunk.
An attention mask that controls which reference and temporal context each video chunk can access.
04 / MEET YOUR NEXT STUDIO
A dedicated local AI workspace for creators who want to explore beyond the ordinary camera feed.
One-time pre-order payment.
05 / THE DETAILS
Core AI inference runs on your GPU machine. Models and software dependencies are downloaded during setup, and you use a browser to open the local interface. A remote GPU machine can also be accessed through a secure tunnel.
For planning, allow a Linux workstation with a modern 8-core CPU, 64 GB RAM, and 150 GB free SSD space. These are setup targets; validated CPU, RAM, and disk minimums have not been published.
The documented configuration has no published 8 GB or 12 GB path. Its lowest reported FP8 file-inference footprint is 14.4 GiB. Lower-resolution or different-model experiments would need their own validation; they are not promised by this site.
The documented path is a CUDA-based Linux GPU environment. Apple GPUs, CPU-only machines, and laptops with integrated graphics are not advertised as supported. Native Windows and WSL2 require additional kernel and camera validation. A lighter computer can access a separate compatible GPU machine through a tunnel.
Yes. The documented workflow accepts source videos, preprocessed video folders, and JSON task files for batch processing. A separate reference image can be used with a source video.
The portraits are AI-generated concept artwork, and the interactive comparison is a style explorer. The transformation demo is the supplied research demonstration with a redesigned Dash presentation. The animated walkthrough and AI presenter guide explain the published training design using adapted diagrams. These educational presentations do not establish new Dash performance measurements. Actual results and performance depend on the model, prompt, input, and hardware.
Pre-order Dash Studio for $599.90 USD through Exnode. The desktop app is in development; a delivery date has not been announced, and a pre-order does not provide an immediate download. GPU hardware and cloud rental are separate. The source workflow is currently described as academic research only; the pre-order does not grant a commercial licence to underlying models.
YOUR NEXT CHARACTER STARTS HERE.