Blog

LTX-2 and LTX-2.3: Specs, Variants, Pricing, and the VFX Tools Nobody Else Ships (Updated August 2026)

Last updated August 7, 2026

LTX-2 and LTX-2.3: Specs, Variants, Pricing, and the VFX Tools Nobody Else Ships (Updated August 2026)

LTX-2 is Lightricks' open-weights audio-video model family; the current generation, LTX-2.3, is a 22B diffusion transformer generating synchronized audio and video at up to native 4K/50fps in ~10-second clips, with portrait support. It is the speed, cost, and openness champion — official API pricing runs $0.08–0.32/s for 2.3 Pro — and uniquely ships VFX utilities: mask-free Clean Plate object removal and SDR-to-HDR conversion.

Updated August 2026

LTX has been the speed, cost, and openness champion of AI video since LTX-Video shipped in November 2024 — and its current generation added something no rival family offers: downloadable VFX utilities, a mask-free object-removal tool (Clean Plate) and an SDR-to-HDR converter, built as lightweight add-ons to the same base model. LTX-2 is Lightricks' open-weights audio-video foundation model family; as of August 2026 the current generation is LTX-2.3, a 22-billion-parameter diffusion transformer generating video with synchronized audio at up to native 4K and 50fps. This page covers the lineage, specs, variants, official pricing, and where LTX-2.3 honestly sits against the top closed models.

Where did LTX-2 come from? The LTX-Video lineage

The LTX line has moved faster than almost any other, and its milestones are unusually well documented because the weights ship in the open.

Date Release What changed
Nov 21, 2024 LTX-Video Lightricks calls it "the first DiT-based video generation model," generating video faster than real time — 768×512 at 24fps initially, open weights (Lightricks/LTX-Video on GitHub)
Mar 2025 0.9.5 Keyframe conditioning; commercial use under an OpenRail-M license
May 2025 0.9.7 13B model with an FP8 build — HD generation in roughly 10 seconds on an H100, per the repo
Jul 2025 0.9.8 Generation length extended up to 60 seconds
Oct 23, 2025 LTX-2 Billed by Lightricks as the first DiT-based audio-video foundation model: synchronized audio + video in one pass, native 4K at up to 50fps, open weights with ComfyUI support (Lightricks/LTX-2 on GitHub)
Early 2026 LTX-2.3 Current 22B generation; technical paper on arXiv, January 2026 (arXiv 2601.03233), with dev and distilled variants on Hugging Face. Lightricks has not published one canonical announcement date for 2.3 — "early 2026" is the honest resolution.

Two lineage facts explain everything else: speed is in the family's DNA, and every generation has shipped open weights — which is why LTX has the deepest community tooling ecosystem of any audio-capable video model.

What are LTX-2.3's specs?

Spec LTX-2.3 (as of August 2026)
Architecture Diffusion transformer (DiT), 22B parameters, joint audio-video generation
Resolution Up to native 4K (Pro tier); 1080p and 1440p tiers below that
Frame rate Up to 50fps
Clip length ~10 seconds is the consistently documented figure; some LTX-2 launch coverage cited 20s, with no definitive official maximum — plan around 10s
Audio Synchronized audio generated natively with the video
Aspect ratios Landscape and portrait (vertical) supported
Inputs Text-to-video, image-to-video, audio-to-video, plus retake and extend operations on the Pro endpoint
Open weights Yes — ltx-2.3-22b-dev (full, trainable) and ltx-2.3-22b-distilled (8-step) on Hugging Face, under the LTX-2 Community License Agreement

Notice which row is soft: resolution, frame rate, and weights are all pinned, but clip length is the one number Lightricks never fixed — plan around 10 seconds and treat anything longer as unconfirmed.

Which LTX variant should you pick?

Inside invideo's model picker, the LTX family appears as one flagship generation model plus two specialized tools — mirroring how Lightricks itself ships the family.

Variant What it is Pick it when
LTX 2.3 Pro High-fidelity generation tier: up to 4K/50fps with synchronized audio, portrait support, and audio-to-video, retake, and extend operations You want the family's best quality for generating new footage, vertical included
LTX 2.3 Clean Plate Video-to-video object removal: deletes people and vehicles and reconstructs the background, no mask required You need to clean existing footage — stray pedestrians, cars, rig gear
LTX Video SDR to HDR Upscale Converts standard-dynamic-range video to HDR You have flat SDR footage that needs HDR color depth for delivery

For local and API users, the other axis is dev versus distilled: the dev checkpoint is the full-quality, trainable model, while the distilled build trades some fidelity for 8-step generation speed.

What makes Clean Plate and SDR-to-HDR unusual?

This part of the LTX story has no equivalent elsewhere. A "clean plate" is a century-old VFX concept — a shot of the scene with nobody in it, used as the background layer for compositing and object removal. Traditionally you film one, or a roto/paint artist builds one frame by frame. Lightricks shipped the task as a model: the LTX-2.3 Clean Plate IC-LoRA turns the base model into a video-to-video remover that deletes people and vehicles from footage and reconstructs the empty background, no mask required — the whole capability a roughly 330MB LoRA riding on the 22B base.

The SDR-to-HDR converter follows the same pattern — shipped as an IC-LoRA and as a hosted "HDR Video (Beta)" endpoint on the first-party API. These two headline a wider IC-LoRA wave through H1 2026: LipDub (dialogue replacement on existing footage), Motion Track, and camera-control LoRAs, most ComfyUI-native on day one. Because the base weights are open, LTX's capabilities grow sideways — small, downloadable, task-specific — in a way closed models structurally can't match.

How much does LTX-2.3 cost?

Lightricks publishes first-party API pricing per second of output (ltx.io/model/api/pricing) — one of the few families where cost planning needs no third-party guesswork:

Tier 1080p 1440p 4K
LTX-2.3 Fast $0.06/s $0.12/s $0.24/s
LTX-2.3 Pro $0.08/s $0.16/s $0.32/s

Specialized endpoints run $0.10–0.40 per second: retake, extend, and audio-to-video at $0.10/s, reframe at $0.10–0.20/s, the HDR beta at $0.20/s (1080p) to $0.40/s (1440p). At $0.08/s, a 10-second 1080p Pro clip with audio costs $0.80 — a fraction of frontier closed-model rates — and self-hosting the open weights removes per-second cost entirely if you have the GPU.

How good is LTX-2.3, honestly?

Its documented strengths: generation speed (at LTX-2's launch Lightricks claimed 18× faster generation than Wan 2.2), native synchronized audio, true 4K at 50fps, portrait support, very low official pricing, open weights, and a VFX-tool ecosystem no rival family has.

The honest caveat: on top-end fidelity and physics, press coverage of the LTX-2 generation has consistently placed it behind the largest closed frontier models — the consensus framing is that LTX competes on speed, cost, and openness rather than beating Veo-class output frame for frame. That is attributed third-party consensus, not a benchmark number. If one hero shot's fidelity is everything, use a frontier model; if you need volume, verticals, audio, or cleanup utilities, LTX-2.3's economics are hard to argue with.

What is LTX-2.3 best used for?

  • High-volume short-form and social video — cheap iteration suits AI video generation workflows where you generate many takes and keep the best.
  • Vertical content — portrait support at the model level, not via cropping, makes it a natural pick for vertical video for Shorts, Reels, and TikTok.
  • Footage cleanup and finishing — Clean Plate's mask-free removal and the SDR-to-HDR pass are AI VFX territory: fixing footage rather than generating it.
  • Dialogue and music-driven clips — audio-to-video on the Pro endpoint generates picture to match a supplied soundtrack or voice line.

How do you prompt LTX-2.3?

The craft that works on other DiT video models applies, with LTX-specific notes:

  • Write shot descriptions, not stories. One subject, one action, one camera move per clip; describe motion in concrete beats.
  • State the audio you want. Audio generates jointly with video, so specifying ambience or dialogue tone ("quiet café hum, soft espresso-machine hiss") gets sound designed for the shot, not generic noise.
  • Declare aspect ratio up front — portrait vs landscape changes composition.
  • Use image-to-video to lock the look, text to direct motion. Anchor a frame you like, then prompt the movement.

Common LTX-2.3 questions

Is LTX-2 open source? The weights are open — LTX-2 and LTX-2.3 checkpoints are downloadable from GitHub and Hugging Face with ComfyUI support — but the license is the LTX-2 Community License Agreement, not an OSI-approved open-source license. "Open weights" is the accurate term.

Which LTX variant should I use? For generating new footage, LTX 2.3 Pro (highest fidelity, 4K/50fps, audio, portrait). For removing people or vehicles from footage, LTX 2.3 Clean Plate. For converting SDR to HDR, the SDR-to-HDR upscale tool. Locally: distilled for speed, dev for maximum quality and trainability.

How long can an LTX-2.3 video be? About 10 seconds per generation is the consistently documented figure as of August 2026; some LTX-2 launch coverage cited 20 seconds, with no definitive official maximum published. The Pro endpoint's extend operation can lengthen a clip beyond one generation.

Can I use LTX-2.3 commercially? The LTX-2 Community License Agreement permits commercial use subject to its conditions — earlier releases (0.9.5+) used OpenRail-M with similar intent. This page won't paraphrase the terms: read the license on the official Hugging Face repo before building a business on the weights.

Does LTX-2.3 generate sound? Yes — synchronized audio and video generate together in one pass, a defining trait of the LTX-2 generation. The Pro endpoint additionally supports audio-to-video: generating picture to match audio you supply.

What hardware does the open model need? The full ltx-2.3-22b-dev checkpoint is demanding — the model card specifies Python 3.12+ and CUDA newer than 12.7, and a 22B DiT wants serious VRAM. The distilled build and the spatial/temporal upscalers exist to make consumer-GPU workflows practical.

Where can you use LTX-2.3 right now?

Self-hosting a 22B model is one path; the shorter one is invideo, the AI video platform that gives serious creatives every major model in one place. LTX 2.3 Pro sits in the agent's video roster next to Veo 3.1, Kling 3.0, Seedance 2.0 and the rest of the 200+ model lineup — and, unusually, so do the family's utilities: LTX 2.3 Clean Plate and the LTX Video SDR to HDR Upscale tool are selectable the same way, so the generate-then-fix loop happens in one timeline instead of across a ComfyUI install. Browse the roster at invideo.io/ai-models, or start from the job: text-to-video, vertical video, or AI VFX.


Version history: LTX-Video Nov 21, 2024 → 0.9.7 (13B) May 2025 → 0.9.8 (60s) Jul 2025 → LTX-2 Oct 23, 2025 → LTX-2.3 (22B) early 2026 → Clean Plate and HDR IC-LoRAs through H1 2026.

Share