Peare AI
Local text-to-video generation with synchronised audio, running entirely on your own machine.
- Role
- ML and creative technology engineer
- Built with
- PyTorch, Diffusers, Stable Video Diffusion, WebSockets, React
- Does
- Creates

Overview
Peare AI is a local-first workstation for generating short videos from text, with an audio track that is generated alongside the picture.
It is built for creators who want control over their results without cloud credits, server queues or subscriptions.
Video diffusion and audio models run directly on the GPU in your own machine.
The problem
Cloud video tools charge per clip, keep you waiting in queues and decide what you are allowed to make.
They also treat sound as an afterthought, so creators end up cutting and syncing audio by hand afterwards.
Peare explores a self-hosted workflow where you write a prompt, guide the result and get a finished clip with sound, all on local hardware.
How it works
- 01
Fitting video diffusion on a consumer GPU
A memory-efficient inference loop with mixed precision and attention offloading keeps temporal diffusion within the limits of consumer graphics cards.
- 02
Keeping sound in sync
Motion detected in the generated frames drives the timing of the generated audio, so hits and sweeps land with the picture.
- 03
A timeline editor
A focused editor lets you guide frames, lock seeds, set camera moves and render straight to disk.
What it does
Fully offline
Runs on your own hardware with no API keys, telemetry or uploads.
Audio that matches the picture
Generated sound is aligned to on-screen motion.
Seed and camera control
Keyframed pans, zooms and locked seeds for consistent results across batches.
Live previews
Frames stream to the editor over WebSockets while they are still being generated.