1. Slug GPU Vector Typography & AE Parity
Why Gitframes
Modern automated video generation is usually constrained by the architectures of general-purpose web browsers: process overhead, non-deterministic DOM layout reflows, and slow screenshot capture. Gitframes treats video composition as software engineering:
Developers generating video programmatically commonly weigh Remotion (React/Chromium) or Hyperframes (Canvas2D/SVG web animation). The matrix below compares the fundamental engineering dimensions.
Capability / Dimension Gitframes Remotion Hyperframes
Underlying Engine Native WebGPU (WGSL compute & render pipelines via Dawn / Metal / Vulkan) Chromium / Puppeteer (React DOM, HTML/CSS layout) Canvas2D / WebGL / SVG (browser or Node Skia)
Rendering Architecture
Direct hardware framebuffer rendering & hardware video encoding (@napi-rs/webcodecs)
Spawns headless Chrome; captures frames via CDP / page.screenshot()
Software or hardware 2D canvas context
Throughput 60–120+ FPS (real-time to faster-than-real-time GPU execution) 5–20 FPS (DOM reflow, IPC, rasterization) 20–40 FPS (CPU draw commands / JS)
Memory Footprint ~200–400 MB per render (zero browser) 1.5–4.0 GB+ per worker (Chromium + V8 DOM heap) ~500 MB–1 GB (Skia/Canvas bindings)
Typography Engine Slug GPU — analytic Bézier evaluation in WGSL, infinite zoom, After Effects selectors Browser DOM text (CSS fonts, rasterized, blurry under 3D transforms) Canvas2D / path text (CPU-rasterized glyphs)
2D VFX & Post-Processing 50+ WebGPU shaders (Curves, Levels, Selective Color, 3D LUT, Film Grain, Halftone, Liquify, PBR Glass, Relight) CSS Filters or custom WebGL canvas wrappers Basic Canvas2D composites and 2D filters
3D Graphics & Depth Native 3D scene graph — LookAt/Turntable camera, multiplane, skinning (OBJ/FBX/glTF), SSAO, PCSS, DoF None built-in (embed Three.js/Fiber inside React DOM) Minimal 2.5D layers; no unified mesh pipeline
Motion Blur & Physics Physical 180° shutter velocity buffers in MRT + closed-form spring kinematics CSS transitions / JS interpolation; synthetic blur hacks Frame interpolation or manual multipass
Audio Engine & DSP
Native audio DSP & procedural SFX (multi-track mixing, beat grids, reactive signals)
<Audio> playback; basic volume curves
Basic static audio playback
Charts & Data Viz
Layer.chart — line, area, bar, scatter, candlestick, pie and donut charts built from native vector nodes, with staggered reveal animations
DOM chart libraries (Recharts, Chart.js)
Custom canvas draw operations
AI & Computer Vision On-device ONNX vision — COCO-80 detection + instance masks (RTMDet-Ins), COCO-17 pose (RTMO), person mattes (Selfie Segmenter); WebGPU tensor conditioning (Canny, depth-to-normals, optical flow, deflicker) External pre-rendered assets; no native GPU tensor conditioning External pre-rendered assets
Headless Verification FrameGrid contact sheets, single-frame snapshots, Skia MSE pixel-invariant assertions Playwright/Puppeteer visual snapshots Manual frame inspection / canvas diffing
Docker / Cloud Portability Compact (~500 MB slim image with native GPU/Vulkan drivers) Heavy (~2–3 GB with Chromium, fonts, X11/Mesa) Moderate container size
Key Features & Capabilities
1. Slug GPU Vector Typography & AE Parity
Traditional text relies on CPU rasterization or low-res SDF atlases that soften under 3D camera sweeps. Gitframes integrates the Slug algorithm (SlugPipeline):
A comprehensive suite of professional image/video shader nodes in nodes/ and packages/webgpu-renderers:
- Calibrated camera rig — LookAt and Turntable cameras (Camera3D) calibrated so z = 0 matches 2D canvas pixel coordinates 1:1.
- 3D layout primitives — Layer3D.cube, carousel, prism, plane, grid with unified depth-buffer testing.
- Zero-dependency model parsers — OBJ, FBX, glTF/GLB, STL, PLY, VOX, 3DS, OFF.
- Skeletal animation & shading — 128-bone Linear Blend Skinning, Blinn-Phong & PBR multi-light shading, PCSS/Poisson contact shadows, SSAO, and optical DoF.
- Physical motion blur — 180° shutter motion blur with per-vertex velocity vectors packed into rg16float MRT buffers.
4. Audio Layers, Procedural SFX & Reactive Signals
- Soundtrack layers — .audio media nodes with frame-exact lifecycle control.
- Procedural SFX — deterministic CPU-synthesized whooshes, impacts, risers, downshifters, and glitches placed on the bar/beat grid (renderSfx, mixSfxInto, softLimit).
- Multi-track mixing — master tracks headlessly with mixAudioTracks and encodeStereoWav.
- Reactive signals — drive transforms, scale, borders, or shader uniforms from tempo signals (Signal.builder) or audio analysis.
5. Animated Charts
Layer.chart builds line, area, bar (grouped or stacked), scatter, candlestick, pie and donut charts. d3 computes the scales, ticks and geometry; every bar, line, slice and label is an ordinary box, path or text node:
Layer.chart( { type: “bar”, width: 900, height: 480, categories: [“Q1”, “Q2”, “Q3”, “Q4”], series: [ { name: “Revenue”, data: [12, 19, 24, 31] }, { name: “Costs”, data: [8, 11, 13, 15] }, ], yAxis: { format: “$,.0f” }, animate: { start: 10, duration: 30 }, }, { position: “absolute”, x: 120, y: 200 }, );
6. On-Device Vision & Tracking
@gitframes/vision runs ONNX models via onnxruntime-node (CPU) or onnxruntime-web (WebGPU) and wires every result into the same reactive signal surface the rest of Gitframes consumes.
Lazy by construction. VisionRunner.create(), comp.withVision(...) and VisionNode.attach(...) perform zero I/O — no downloads, no sessions, no file probes. A model is fetched the first time a task actually runs. To warm up ahead of time, call await runner.preload(["detect", "pose"]) (or await vision.ready() on an attached node).
Every model is Apache-2.0, pinned to an immutable Hugging Face revision, and verified by SHA-256 after download.
- OpenPose-style skeleton textures — rasterize COCO-17 keypoints into a VRAM conditioning texture (PoseSkeletonRenderer).
- GPU segmentation texture pool — reusable silhouette textures (SegmentationTexturePool).
Temporal tracking & analysis
- Multi-object tracker (TemporalObjectTracker) assigns stable trackIds via IoU association, with configurable minHits, positionSmoothing, and velocity-based coasting for up to maxMissedFrames (default 15) so a transient miss holds the track instead of flashing.
- Pose↔track matching (pose-track-matcher) binds keypoints to the right track by id, then by spatial IoU fallback.
- One-shot sequence analysis — comp.analyzeVisionSequence(src, { tasks, categories }) decodes frames through the mediabunny pipeline, tracks them, and returns a zod-serializable report (per-track frame ranges, mean speed, sampled center paths, per-class presence/confidence, mean mask coverage, model download bytes/timing) (analyzeSequence).
Reactive vision signals
Every tracked entity is exposed as reactive ProgrammaticSignals that animate layers and shader uniforms:
Project normalized landmarks to screen space with a configurable camera FOV, then bind any node to a track or landmark (SpatialLandmarkTransformer, spatial-pin):
- Subject Sandwich — comp.addSubjectSandwich({ source, behind, feather, fit }) cuts the foreground subject out and places typography/graphics behind them.
- Smart Reframing — comp.addSmartFraming({ source, target, targetAspect, damping, leadHeadroom }) auto-crops 16:9 → 9:16 while tracking target.
- Subject Outline — comp.addSubjectOutline(vision.segmentation.subject, { source, color, width, blur }) strokes the segmented boundary as an audio-reactive contour glow.
- Tracked Region Blur — layer.blurRegion(track, { strength }) blurs faces, plates, or any detected class.
- Node modes — passthrough, mask, matte, crop, skeleton, boxes, tracking; pick the cutout alpha with matteSource: “instance” | “selfie”, and optionally keyBackground to grow the subject into connected foreground.
Agent-first DX
- Runtime config is zod-validated and available from a zod-only entry (@gitframes/vision/schemas) so the hot path stays zod-free. Unknown or removed options are rejected, not silently ignored.
- vision.summary(frame) returns a deterministic, serializable snapshot (objects, classes, masks) safe to call inside a frame hook.
- Clear failures — a model that is the wrong size, fails its checksum, or lacks an expected output raises an error naming the model and its source.
- Browser entry — @gitframes/vision/web re-exports the engine plus createWebGPUProvider() / hasWebGPU(); onnxruntime-web is an optional lazy peer.
7. Headless Conformance & FrameGrid Testing
- Pixel-sampling invariant assertions — test compositions in Vitest with skia-canvas to verify shader math, font coverage, and Mean Squared Error (MSE) temporal deltas.
- FrameGrid contact sheets — comp.renderFrameGrid(…) outputs sequential-frame contact sheets for instant review of easing, kinetic type, and transitions.
8. Live Preview in the Browser
- Runs your composition, not a video — startPreview({ entry, export }) serves a localhost WebGPU player that loads the composition’s own module and renders every frame live in the browser. Nothing is streamed: the server only hands over the bundle, the project’s assets, and the soundtrack mixed by the export engine.
- Timeline, waveform & frame stepping — play/pause, scrub, step frame by frame, and read resolution, FPS, duration, and audio status at a glance.
- One stable URL per project — the port is derived from the working directory, so re-running the preview replaces the running server and any open tab reloads into the new version by itself. Close the tab and the server shuts down about five seconds later.
- Shown where you are — startPreview serves the page and returns its URL instead of opening a browser, so an agent can show it in its own pane (Claude Code, Codex); pass open: true to open the system browser.
import { startPreview } from "gitframes";
const session = await startPreview(
{ entry: new URL("./film.ts", import.meta.url), export: "buildFilm" },
{ title: "gitframes film" },
);
console.log(`Preview at ${session.url}`);
await session.closed; // serves until its tab closes or a newer preview takes over
gitframes/ ├── packages/ │ ├── gitframes/ # Unified SDK (Composition, Layer, LayerAnimation, Signal, effects) │ ├── core/ # Core AST, Effect base class, VirtualMediaData, vision types │ ├── compositions/ # Layout engine, Flex/Box AST compiler, timeline evaluator │ ├── webgpu-renderers/ # WGSL shaders, Slug text engine, 3D renderer, camera, lights, materials │ ├── tensor-webgpu/ # WebGPU compute pipelines (Canny, depth-to-normals, flow, deflicker, landmarks) │ ├── vision/ # ONNX vision engine: detect, segment, pose, matte, tracking, signals │ ├── renderer/ # Headless Node.js WebGPU renderer via Dawn, WebCodecs, skia-canvas │ ├── renderers/ # Higher-level render orchestration │ ├── media/ # Media decoding / encoding adapters │ ├── node-sdk/ # Node renderer contracts and result schemas │ ├── server-utils/ # Server infrastructure, storage, asset caches │ └── client-utils/ # Shared browser utilities ├── nodes/ # 58+ specialized domain nodes (VFX, audio, layout, node-vision) ├── apps/ │ └── renderer-service/ # Production HTTP / gRPC rendering microservice container ├── examples/ # Reference compositions and films ├── plugins/gitframes/ # Agent plugin: skills only (setup, compose, effects, render) └── scripts/ # Build, release, and plugin validation tooling
Quickstart Guide
Installation
pnpm add gitframes
Requirements: Node.js ≥ 22. Gitframes uses native GPU acceleration via Dawn / WebGPU or Vulkan.
import { Composition, Layer, LayerAnimation } from "gitframes";
// 1. Initialize a 1080p60 composition
const comp = new Composition({
width: 1920,
height: 1080,
fps: 60,
durationFrames: 180, // 3 seconds
backgroundColor: "#090a0f",
fonts: ["assets/fonts/Inter.ttf", "assets/fonts/SpaceGrotesk.ttf"],
});
// 2. Define physical snap-overshoot animations
const cardEntrance = LayerAnimation.create()
.fadeIn(0, 20, "power2.out")
.fromTo("y", 60, 0, { start: 0, end: 35, ease: "back.out(1.5)" })
.fromTo("scale", 0.92, 1.0, { start: 0, end: 35, ease: "back.out(1.2)" });
// 3. Assemble a responsive flex-layout card
const heroCard = Layer.box({
width: 720,
height: 380,
background: "#141721",
borderRadius: 24,
borderColor: "#262b3d",
borderWidth: 1.5,
padding: 32,
children: [
Layer.flex({
dir: "column",
gap: 16,
children: [
Layer.text("GITFRAMES ENGINE", {
fontSize: 16,
fontWeight: 700,
fill: "#6366f1",
letterSpacing: 2.0,
}),
Layer.text("Next-Gen WebGPU Motion", {
fontSize: 48,
fontWeight: 700,
fill: "#f8fafc",
fontFamily: "SpaceGrotesk",
}),
Layer.text("Direct hardware video composition without headless browser overhead.", {
fontSize: 20,
fill: "#94a3b8",
lineHeight: 28,
}),
],
}),
],
}).animate(cardEntrance);
comp.add(heroCard);
2. Unified 3D Scene with Camera & 3D Model
import { Composition, Layer, Layer3D, CameraAnimation, Light } from "gitframes";
const comp = new Composition({ width: 1920, height: 1080, fps: 60, durationFrames: 300 });
// 1. LookAt 3D camera with a continuous orbit
const cameraAnim = CameraAnimation.camera().orbit({
azimuth: { from: -30, to: 30 },
elevation: { from: 15, to: 15 },
radius: { to: 1200 },
start: 0,
end: 300,
});
comp.add(
Layer.camera({ x: 960, y: 540, z: -1000, targetX: 960, targetY: 540, targetZ: 0 }).animate(cameraAnim)
);
// 2. Studio lighting
comp.add(Light.ambient("#ffffff", 0.4));
comp.add(Light.directional({ color: "#e0e7ff", intensity: 1.2, x: 500, y: -800, z: -600 }));
// 3. 3D model with skeletal animation
comp.add(
Layer.glb("assets/models/character.glb", {
x: 960,
y: 640,
z: 0,
scale: 2.5,
material: "lit",
loop: true,
})
);
// 4. 3D prism layout carousel
comp.add(
Layer3D.carousel({
radius: 400,
items: [
Layer.box({ width: 280, height: 180, background: "#1e293b", borderRadius: 16 }),
Layer.box({ width: 280, height: 180, background: "#334155", borderRadius: 16 }),
Layer.box({ width: 280, height: 180, background: "#0f172a", borderRadius: 16 }),
],
})
);
3. Audio Soundtrack, Procedural SFX & Reactive Signals
import { Composition, Layer, LayerAnimation, Signal, renderSfx, mixSfxInto, softLimit } from "gitframes";
const comp = new Composition({ width: 1920, height: 1080, fps: 60 });
const totalFrames = 240;
// 1. Soundtrack layer
comp.addAudio(Layer.audio("assets/score.mp3", { volume: 0.9, durationFrames: totalFrames }));
// 2. Frame-accurate procedural SFX on the beat grid
const bed: [Float32Array, Float32Array] = [
new Float32Array(Math.ceil((totalFrames / 60) * 48000)),
new Float32Array(Math.ceil((totalFrames / 60) * 48000)),
];
mixSfxInto(bed, [
renderSfx({ type: "whoosh", atBar: 0.79, volume: 0.5 }, { sampleRate: 48000, secondsPerBar: 2.0, seed: 1 }),
renderSfx({ type: "impact", atBar: 1.0, volume: 0.8 }, { sampleRate: 48000, secondsPerBar: 2.0, seed: 2 }),
]);
softLimit(bed);
// 3. Tempo signal (120 BPM = 2 Hz)
const beatPulse = Signal.builder({ type: "sawtooth", frequency: 2, amplitude: 0.08, offset: 1.0 });
// 4. Bind it to visuals
const reactiveCard = Layer.box({ width: 400, height: 250, background: "#1c202e", borderRadius: 20 })
.animate(
LayerAnimation.create()
.signal("scale", beatPulse, { multiplier: 1.0, offset: 0.0 })
.fromTo("opacity", 0, 1, { start: 0, end: 15, ease: "power2.out" }),
);
comp.add(reactiveCard);
4. Chained WebGPU Post-Processing VFX
import { Composition, FilmGrain, Vignette, ColorBalance } from "gitframes";
const comp = new Composition({ width: 1920, height: 1080, fps: 60 });
// Whole-composition cinematic grade + film emulsion
comp.apply(new Vignette({ strength: 0.28, radius: 0.85 }));
comp.apply(new FilmGrain({ strength: 0.06, size: 1.5, animated: true }));
comp.apply(
new ColorBalance({
shadows: { cyanRed: 0, magentaGreen: 2, yellowBlue: 6 },
highlights: { cyanRed: 4, magentaGreen: 1, yellowBlue: -2 },
}),
);
5. Vision: Pin, Matte & Reframe
import { Composition, Layer, Vignette } from "gitframes";
const comp = new Composition({ width: 1920, height: 1080, fps: 30 });
// Run vision on the whole composition. Models download lazily on first use.
const vision = comp.withVision({
enableDetection: true,
enableSegmentation: true,
enablePose: true,
variant: "s",
confidence: 0.35,
});
// Pin a caption to the primary tracked subject (smoothing + auto-hide when lost)
comp.add(
Layer.text("SUBJECT 01", { fontSize: 40, fill: "#f8fafc" }).pinToObject(
vision.objects.primary,
{ anchor: "topCenter", offsetY: -48, smoothFrames: 5, hideWhenLost: true },
),
);
// Drive a shader uniform from a reactive signal — here, subject mask coverage
comp.add(
Layer.box({ width: 1920, height: 1080, background: "#000000" }).withEffect(
new Vignette({ strength: vision.segmentation.subject.coverage, radius: 0.9 }),
),
);
// Or use the one-liners for the common editorial moves:
// comp.addSubjectSandwich({ source: "assets/dancer.mp4", behind: [headline], feather: 4 });
// comp.addSmartFraming({ source: "assets/action.mp4", target: vision.objects.primary, targetAspect: 9 / 16 });
// comp.addSubjectOutline(vision.segmentation.subject, { source: "assets/character.mp4", color: "#FF5A1F", width: 6 });
// Inspect a source before authoring: one-shot, ffmpeg-free, zod-serializable report
const report = await comp.analyzeVisionSequence("assets/street.mp4", {
tasks: ["detect", "pose"],
categories: ["person"],
});
console.log(report.tracks.map((t) => `${t.category}#${t.trackId} ${t.frames.join("–")}`));
Standalone runner (no composition):
import { VisionRunner } from “@gitframes/vision”;
const runner = VisionRunner.create({ variant: “s”, confidence: 0.3 }); // zero I/O const frame = { data: rgba, width: 1920, height: 1080 }; const boxes = await runner.detect(frame); // downloads RTMDet-Ins on first call const { masks } = await runner.segment(frame); // same forward pass, no second inference const { people } = await runner.pose(frame); // RTMO, COCO-17 keypoints runner.close();
In the browser (WebGPU EP):
import { VisionRunner, createWebGPUProvider, hasWebGPU } from “@gitframes/vision/web”;
if (hasWebGPU()) { const runner = VisionRunner.create({ provider: createWebGPUProvider() }); }
6. Headless Video & FrameGrid Rendering
import { buildMyComposition } from "./my-composition.js";
const comp = await buildMyComposition();
// 1. Single frame to a PNG buffer for visual inspection
const frameBuffer = await comp.renderFrame({ frame: 45 });
// 2. Contact-sheet grid of 12 sequential frames
const gridBuffer = await comp.renderFrameGrid({
startFrame: 0,
endFrame: 120,
stepFrames: 10,
cellWidth: 320,
showLabels: true,
});
// 3. Final hardware-encoded MP4 with mixed audio
const { filePath } = await comp.renderVideo({
outputPath: "output/final-product-film.mp4",
quality: "high",
concurrency: 4,
});
console.log(`Video rendered successfully to: ${filePath}`);
Engineering Doctrines & Best Practices
- Design tokens & theme contracts — define a centralized THEME for colors, type, radii, and spacing. Never hardcode magic hex values or ad-hoc margins.
- WebGPU premultiplied-alpha invariant — fragment shaders outputting premultiplied alpha (color * opacity * alpha) must use srcFactor: “one” in their blend state ({ srcFactor: “one”, dstFactor: “one-minus-src-alpha”, operation: “add” }). Never use srcFactor: “src-alpha” for premultiplied output — squaring alpha darkens fades into murky gray.
- Carrier match cuts — carry a visual element (badge, card, cursor, container) across scene boundaries with continuous velocity and position to avoid jarring cuts.
- Physical easing vocabulary — back.out(1.4–1.7) for snap-overshoot entrances, spring / expo.out for decelerating motion, power2.in for exits. Reserve linear for infinite spinners and time counters.
- Headless invariant verification — verify shader transforms, glyph coverage, and temporal MSE deltas with skia-canvas pixel sampling in Vitest before shipping.
Agent Skills & Plugins
Gitframes ships agent skills that teach Claude, Codex, and other coding agents how to write, render, and check compositions. The plugin (gitframes) is listed in Anthropic’s official plugin directory and contains only skills — no MCP servers, hooks, or commands. Every other agent gets the same skills through the skills CLI.
Once installed, skills load automatically when a task matches (e.g. “add a film-grain pass to this scene” or “render a frame grid of intro.ts”).
The plugin is instructions only. It bundles no executables, MCP servers, hooks, or package launchers, and it sends no data anywhere. The skills tell your agent to add the gitframes npm package to your project and how to use it. When that code uses on-device vision, the SDK downloads the pinned model weights from Hugging Face on first use (see On-Device Vision). Nothing else leaves your machine.
/plugin install gitframes
Or from your shell:
claude plugin install gitframes@claude-plugins-official
It installs from Anthropic’s official marketplace, which Claude Code adds for you, so there is no marketplace step, and plugins from it update automatically. Afterwards, restart Claude Code or run /reload-plugins. /plugin commands need an interactive claude terminal; in the desktop app’s Code tab, use the shell form or + > Plugins > Add plugin and pick Gitframes.
Add --scope project to the shell form to record the plugin in .claude/settings.json for the whole team.
Enable it for everyone in your repo. Commit this to .claude/settings.json; Claude Code prompts teammates to install it when they trust the folder:
{ “enabledPlugins”: { “gitframes@claude-plugins-official”: true } }
Straight from this repository (tracks main instead of the directory release):
/plugin marketplace add gatewai-dev/gitframes /plugin install gitframes@gitframes-plugins
Codex, Cursor, Hermes, and other agents
The skills CLI installs the skills into any of 70+ agents, including Codex, Cursor, Hermes, Gemini CLI, GitHub Copilot, Windsurf, OpenCode, and Goose:
npx skills add gatewai-dev/gitframes
It detects the agents on your machine and asks where to install. To choose them yourself, pass -a once per agent, add -g to install for your user instead of this project, and -y to skip the prompts:
