Self-Host an AI Video Generator: Open Models, GPU Costs and Real Cost per Second (2026)
- Hosted video APIs charge $0.03 to $0.70 per generated second in September 2026, and OpenAI's Sora 2 API shuts on 24 September 2026. Self-hosting buys control of the model, the data and the roadmap first; a lower bill arrives only with volume.
- Read the licence before the leaderboard. The top open model, MiniMax H3, is not licensed in the US, EU, UK or South Korea; LTX-2.5 matches Veo 3.1 Lite in silent-video rankings and is free below $10M revenue; Wan 2.2 is Apache 2.0 but about 100 Elo behind.
- Stock Wan 2.2 on a rented H100 costs about $0.50 per generated second at 40% utilisation, more than Veo 3.1 with audio. A 4-step distilled build brings the marginal cost under a cent, and my modelled break-even against Veo 3.1 Fast is about 87,000 seconds of video a month.
Templates, brand kit, a script-to-shot planner, a 480p preview and a 720p final render. This is the part users pay for.
where the value isAuth, per-account quotas, an idempotency key per request, a policy classifier on the prompt and a consent check on any uploaded face.
fails: abuse at scaleA durable queue with an interactive lane and a bulk lane. Workers scale on queue depth and the age of the oldest job, never on CPU.
fails: GPU starvation at peakWan 2.2 or LTX-2.5 with warm weights, a pinned distilled checkpoint and brand LoRAs, driven by ComfyUI or plain diffusers.
22.5 s per 5 s clip on an RTX 5090Encode, sample frames through an image safety classifier, embed an invisible watermark and sign a C2PA manifest.
EU marking duty from 2 Aug 2026MP4 files in object storage behind signed CDN URLs, plus one ledger row per job with seed, versions, GPU-seconds and cost.
R2 egress: $0What does it cost to add AI video generation with an API in 2026?
A feature that turns product pages into short videos pays $0.03 to $0.70 per second of finished video. On Vertex AI, Veo 3.1 Lite is $0.03 a second for silent 720p and Veo 3.1 Fast $0.10 with audio. Runway Gen-4.5 is $0.12, Veo 3.1 with audio $0.40, and Sora 2 Pro at 1080p $0.70 until it shuts.
Start with the feature, because the feature sets the volume. The common shape is short-form video inside a product people already pay for: a merchant pastes a product page and gets five vertical ads, a property listing becomes a walkthrough, a help article becomes a 20-second explainer. Take 2,000 merchants rendering 20 six-second clips a month each. That is 240,000 seconds of video: $24,000 a month on Veo 3.1 Fast with audio, $7,200 on Veo 3.1 Lite without it, and $96,000 on Veo 3.1 with audio. The quality tier moves the bill thirteen-fold before any architecture decision is made.
Two other costs push teams toward running their own model. The first is data: every prompt, product photo and brand asset goes to the vendor, and some retail, healthcare and enterprise buyers will not sign that off. The second is that the roadmap belongs to the vendor. OpenAI announced on 24 March 2026 that the Sora 2 models and its Videos API shut down on 24 September 2026, with no replacement listed on its deprecations page. If your templates were tuned to one model's look, a retirement notice means re-tuning all of them on someone else's schedule.
A note on what follows. This is a reference design and a cost model, not a case study: I have not run this exact stack for a client. Every price is a published list price checked on 23 September 2026 and linked, every speed figure is a published benchmark with its hardware named, and every dollar figure I derive shows its arithmetic. Where I estimate, I say so.
Which open-weights video models can you self-host and legally ship?
In September 2026, two families are both good and usable by a US or EU company: LTX-2.5 (22B parameters, video plus audio, free below $10M annual revenue) and Wan 2.2 (Apache 2.0, silent). The arena leader, MiniMax H3, and Tencent's HunyuanVideo 1.5 carry licences that exclude major markets.
The licence is now the first filter and the leaderboard the second. Alibaba shipped Wan 2.5, 2.6, 2.7 and 3.0 as closed models, so Wan 2.2 from July 2025 is still its newest open flagship, while Wan 3.0 sits second on the Artificial Analysis text-to-video arena with no weights at all. MiniMax H3, released in August 2026, is the best open model on that arena at 1,220 Elo with audio, well above Veo 3.1 at 1,088. Its licence excludes the United States, the EU, the UK and South Korea, forbids using its outputs there, requires written authorisation above $20 million of yearly revenue and makes you display the model's name in your interface.
That leaves two defaults. LTX-2.5 is the quality pick: in the arena's silent-video ranking it scores 1,217 against 1,216 for Veo 3.1 Lite and 1,213 for Veo 3.1, it generates synchronised audio, and the LTX-2.x Community License is royalty-free until your company reaches $10 million in annual revenue, after which production use needs a paid agreement with Lightricks. Wan 2.2 is the freedom pick: Apache 2.0 with no thresholds and a large ecosystem of LoRAs, ComfyUI workflows and distilled checkpoints, at the price of silence and a score about 100 Elo below Veo 3.1 Lite (1,116 against 1,216).
The rest are situational. HunyuanVideo 1.5 runs in 14 GB with offloading and has a fast 480p distilled path, but Tencent's licence does not apply in the EU, UK or South Korea and bars use of its outputs there, which rules it out for most consumer products with European users. Robbyant's LingBot-Video, an Apache 2.0 mixture-of-experts model from July 2026, targets robotics and renders at 480 by 832. Mochi 1 and CogVideoX are 2024 models that still run but lose most arena match-ups, and CogVideoX-5B needs registration for commercial use and caps free use at one million visits a month.
| Model | Released | Size | Licence and commercial terms | Audio | Published VRAM floor | Arena Elo (silent / with audio) |
|---|---|---|---|---|---|---|
| MiniMax H3 | Aug 2026 | 33B | MiniMax H3 Community: excludes US, EU, UK and South Korea; authorisation above $20M revenue | Yes | Not published | 1,302 / 1,220 |
| LTX-2.5 | Aug 2026 | 22B | LTX-2.x Community: free below $10M annual revenue, paid above | Yes | 16 GB (vendor figure) | 1,217 / 1,055 |
| LTX-2.3 | Mar 2026 | 22B | LTX-2 Community: same $10M threshold | Yes | Not published | 1,121 / 975 |
| Wan 2.2 A14B | Jul 2025 | 27B MoE, 14B active | Apache 2.0, no thresholds | No | 80 GB stock at 720p; 24 GB with distilled FP8 | 1,116 / not ranked |
| Wan 2.2 TI2V-5B | Jul 2025 | 5B | Apache 2.0 | No | 24 GB (RTX 4090) | Not ranked |
| HunyuanVideo 1.5 | Nov 2025 | 8.3B | Tencent Hunyuan Community: not licensed in EU, UK or South Korea | No | 14 GB with offloading | 1,020 / not ranked |
| LingBot-Video | Jul 2026 | 30B MoE | Apache 2.0 | No | Not published | Not ranked |
| Mochi 1 | Oct 2024 | 10B | Apache 2.0 | No | About 60 GB; under 20 GB in ComfyUI | 1,000 / not ranked |
| CogVideoX-5B | Aug 2024 | 5B | CogVideoX License: registration, free to 1M visits a month | No | 5 GB with diffusers offloading | 805 / not ranked |
What GPU do you need for AI video generation, and what does an hour cost?
Stock Wan 2.2 at 720p needs an 80 GB card such as an H100 or A100; distilled FP8 builds run in 24 GB, and the fastest 4-step NVFP4 build needs a Blackwell card. On demand, an H100 costs $3.49 an hour on RunPod and $12.29 on Azure; an RTX PRO 6000 is $2.09.
The same H100 costs 3.5 times as much on Azure as on RunPod. The hyperscalers price on-demand GPUs for customers who will sign commitments; the GPU clouds sell on-demand capacity as the product. The practical pattern is to prototype on RunPod, Lambda or Modal and move to a hyperscaler only when credits or a committed-use discount change the arithmetic. Credits change it a lot: I have secured $300K of them (Microsoft for Startups $200K, AWS Activate $100K), and how to get cloud credits covers the route. Model your own fleet in the cloud cost calculator.
Buying is cheaper only if the cards stay busy. NVIDIA's marketplace price for an RTX PRO 6000 rose to $16,000 in August 2026, from $8,565 at launch, and an H100 costs $27,000 to $40,000 depending on vendor. Two RTX PRO 6000s plus a host (about $8,000, my estimate) is $40,000, or $1,111 a month over 36 months. Two 600 W cards and about 300 W for the host draw 1.5 kW; at the US commercial average of 14.19 cents per kWh in June 2026 that is about $155 a month at full load. Renting the same two cards on RunPod is $3,051 a month, so the purchase pays back in about 14 months of full-time use, before colocation, spares and the person who fixes it at 2am.
One licensing trap: NVIDIA's GeForce driver licence (section 2.8) says the software is not licensed for datacenter deployment. An RTX 5090 on your desk is a fine development machine, and renting one from a cloud is that provider's licensing question, but a rack of 5090s in your own server room is not a production plan. The RTX PRO 6000 is the Blackwell card built for that job.
| GPU (memory) | RunPod | Lambda | Modal | CoreWeave | AWS | Google Cloud | Azure |
|---|---|---|---|---|---|---|---|
| L40S (48 GB) | $1.09 | Not listed | $1.95 | $2.25 | $1.86 (g6e.xlarge) | Not quoted | Not quoted |
| A100 (80 GB) | $1.59 | $2.79 | $2.50 | $2.70 | Not quoted | $5.07 (a2-ultragpu-1g) | Not quoted |
| RTX PRO 6000 (96 GB) | $2.09 | Not listed | $3.03 | $2.50 | Not quoted | $4.50 (g4-standard-48) | Not quoted |
| H100 (80 GB) | $3.49 | $3.99 | $3.95 | $6.16 | $6.88 (p5.4xlarge) | $11.06 (a3-highgpu-8g) | $12.29 (ND H100 v5) |
| H200 (141 GB) | $4.59 | Not listed | $4.54 | $6.31 | $7.91 (p5en.48xlarge) | $10.60 (a3-ultragpu-8g) | Not quoted |
| B200 (180 GB) | $6.79 | $6.69 | $6.25 | $8.60 | $14.24 (p6-b200.48xlarge) | No on-demand; $8.06 flex-start | Not quoted |
How many seconds of video does one GPU produce per minute?
Anywhere from 0.1 to 44 seconds, and the step count decides it. Stock Wan 2.2 A14B takes 1,041.5 seconds on one H100 for a five-second 720p clip: 0.29 seconds of video per GPU-minute. A 4-step NVFP4 build of the same model takes 22.5 seconds on an RTX 5090: 13.5 seconds per GPU-minute.
Two numbers explain the spread. Stock Wan 2.2 A14B runs 40 denoising steps, each with a conditional and an unconditional pass for classifier-free guidance, so the large transformer runs 80 times per clip. A 4-step distilled student that needs no guidance runs it four times. That 20-fold cut in passes, plus NVFP4 weights and sparse attention, is how LightX2V reports going from 2,668 seconds to 22.5 seconds on the same RTX 5090. The quality claim is theirs; run your own side-by-side on your prompts before you believe it.
Parallelism buys latency, not cost. Morphic's optimised run finishes a 720p image-to-video clip in 109.8 seconds on eight H100s, which is 878 GPU-seconds, only a little better than the 1,055.9 seconds the Wan team reports for the same job on one H100. Splitting a job across GPUs is worth it when a user is waiting and the clip is long; for batch ad rendering, one clip per GPU keeps the cost per second lowest.
Treat every figure in the table as a best case: warm weights, one clip at a time, the publisher's own hardware and software. The LTX-2.5 number was measured on two GB200 superchips, which few teams can rent by the hour, and nobody has published an RTX PRO 6000 figure for it yet. Budget a day to benchmark your chosen model on the exact card you will rent, at your resolution and clip length, before you size a fleet on anyone's blog post, including this one.
| Configuration | Hardware | Time per clip | Video seconds per GPU-minute | Source |
|---|---|---|---|---|
| Wan 2.2 A14B, 720p, 40 steps | 1x H100 | 1,041.5 s for 5 s | 0.29 | Wan team README |
| Wan 2.2 A14B, 720p, 40 steps | 1x A100 80GB | 2,735.7 s for 5 s | 0.11 | Wan team README |
| Wan 2.2 A14B I2V, 720p, 40 steps, FA3, MagCache, compile | 8x H100 | 109.8 s for 5 s | 0.35 | Morphic |
| LightWan2.2 A14B, 720p, 4 steps, NVFP4, sparse attention | 1x RTX 5090 | 22.5 s for 5 s | 13.5 | LightX2V |
| FastWan2.2 TI2V-5B, 720p, few steps | 1x H200 | 16 s for 5 s | 18.9 | Hao AI Lab |
| HunyuanVideo 1.5 I2V, 480p, 8 to 12 steps | 1x RTX 4090 | 75 s for 5 s | 4.0 | Tencent README |
| LTX-2.5 distilled, 720p | 2x GB200 | 6.8 s for 10 s | 44.1 | Lightricks |
What does one generated second cost, and when does self-hosting break even?
The marginal cost is GPU price per second times GPU-seconds per video-second, divided by utilisation. At 40% utilisation that is about $0.50 for stock Wan 2.2 on an H100 and about $0.0065 for the 4-step build on an RTX PRO 6000. Against Veo 3.1 Fast, my modelled break-even is about 87,000 seconds a month.
Work one row through. RunPod charges $2.09 an hour for an RTX PRO 6000, which is $0.00058 a second. If a five-second clip takes 22.5 GPU-seconds (an estimate: the published figure is for an RTX 5090, a smaller Blackwell card), a clip costs $0.0131 and a generated second $0.0026. GPUs are rarely busy all day. I assume 40% average utilisation for traffic that peaks in office hours on a fleet sized for the peak, so divide by 0.4: $0.0065 a second. Utilisation is the whole economic question here, the same way it is for LLM inference cost.
The marginal cost is not the bill. The bill is a floor you pay whether or not anyone renders a clip: two GPUs always on, so a crash or a deploy does not take the feature down, plus the people who run it. Modelled at $1,000 per engineer-day (a $250,000 fully loaded salary over 250 working days): two RTX PRO 6000s on RunPod at $3,051 a month, four engineer-days of upkeep at $4,000, and an eight-week build of $40,000 amortised over 24 months at $1,667. That is $8,718 a month before the first clip, and the two cards can produce about 473,000 seconds of video a month at 40% utilisation.
Divide the floor by the API price you would otherwise pay. Against Veo 3.1 Fast at $0.10 a second, self-hosting wins above about 87,000 seconds a month, roughly 17,400 five-second clips. Against Veo 3.1 Lite without audio at $0.03, it needs about 290,000 seconds. For the 2,000-merchant example at 240,000 seconds, the self-hosted stack costs $8,718 against $24,000 on Veo 3.1 Fast, and $1,518 more than Veo 3.1 Lite. The same arithmetic governed the voice AI stack I moved from about 10 cents a minute to 2.5 cents on a custom LiveKit build: a lower unit price only pays back once the volume is there. Price your own version in the AI product cost estimator.
| Setup | GPU price per hour | Cost per clip | Per video second at 100% busy | At 40% utilisation |
|---|---|---|---|---|
| Wan 2.2 A14B stock, 1x H100 (RunPod) | $3.49 | $1.01 per 5 s | $0.199 | $0.499 |
| Wan 2.2 A14B stock, 1x A100 80GB (RunPod) | $1.59 | $1.21 per 5 s | $0.239 | $0.597 |
| LightWan2.2 4-step, 1x RTX 5090 (RunPod) | $0.99 | $0.0062 per 5 s | $0.0012 | $0.0031 |
| LightWan2.2 4-step, 1x RTX PRO 6000 (RunPod, 5090 speed assumed) | $2.09 | $0.0131 per 5 s | $0.0026 | $0.0065 |
| FastWan2.2 5B, 1x H200 (RunPod) | $4.59 | $0.0204 per 5 s | $0.0040 | $0.0101 |
| HunyuanVideo 1.5 480p, 1x RTX 4090 (RunPod) | $0.74 | $0.0154 per 5 s | $0.0031 | $0.0076 |
| LTX-2.5 distilled, 2x GB200 (CoreWeave) | $10.50 per GPU | $0.0397 per 10 s | $0.0040 | $0.0099 |
What does the serving architecture for self-hosted video generation look like?
Nine parts in a line: an API gateway, a prompt policy check, a durable job queue, warm GPU workers, post-processing, output moderation, a C2PA signer, object storage behind a CDN, and a job ledger. Generation is asynchronous: the app submits a job, gets an id back at once, and hears about completion by webhook.
The queue carries more design than it looks. Give it two lanes: an interactive lane for 480p previews that users wait on, and a bulk lane for final 720p renders and batch jobs. Put an idempotency key on every submission so a double-click does not buy two clips, and set the visibility timeout longer than your slowest generation so a slow worker does not cause the same clip to be rendered twice. Scale workers on queue depth and the age of the oldest job, never on CPU, because a GPU worker at full GPU load and 5% CPU looks idle to a default autoscaler.
On the workers, pick the driver by who owns the pipeline. If designers build and change workflows, run ComfyUI headless and submit jobs to its server API: POST /prompt to queue, progress over the /ws websocket. If engineers own it, a plain Python worker on diffusers or LightX2V is easier to test, version and profile. Either way, keep one model per worker process with weights loaded at start, because loading tens of gigabytes per request would erase the speed gains above, and pin every version: checkpoint hash, LoRA hash, sampler, step count and seed. Brand LoRAs let one warm base model serve many customers' looks.
Everything after the GPU is ordinary backend work. Encode to H.264 on the card's NVENC encoders where it has them: per NVIDIA's codec support matrix, an RTX PRO 6000 Server Edition has four and an L40S three, while an H100 or A100 has none, so budget CPU time for FFmpeg on those. Store MP4 files in object storage (Cloudflare R2 is $0.015 per GB-month with free egress), hand the app signed URLs, and delete abandoned drafts on a lifecycle rule. Write one ledger row per job with prompt, seed, versions, GPU-seconds, cost and moderation verdicts. When a customer disputes a clip three weeks later, that row is the answer. The general layering is in the AI product architecture guide.

- 1Submitsynchronous
The app posts prompt, reference image and brand kit with an idempotency key. The gateway checks quota and returns a job id at once.
- 2Check the promptbefore any GPU time
A policy classifier screens the text; a face detector routes uploads of real people to a consent check.
- 3Queueidempotent, at-least-once
The job lands in the interactive or bulk lane. Workers pull by priority; the scheduler scales on queue depth.
- 4Generate22.5 s per 5 s clip on an RTX 5090 (published)
A warm worker runs the pinned distilled checkpoint with the customer's LoRA and a recorded seed.
- 5Check and signblocked clips never reach storage
Encode, sample one frame per second through an image classifier, embed a watermark, sign the C2PA manifest.
- 6Deliverasynchronous to the user
Upload to object storage, write the ledger row, fire the webhook with a signed URL.
How do you handle moderation, watermarking and C2PA provenance for AI video?
Screen the prompt and any uploaded image before the GPU runs, classify sampled frames after, then embed an invisible watermark and sign a C2PA manifest that marks the file as AI-generated. Since 2 August 2026, EU AI Act Article 50(2) requires machine-readable marking of AI-generated video, and California's AI Transparency Act took effect the same day.
Moderation is cheaper before generation than after, because a blocked prompt costs a fraction of a render and a blocked clip costs the whole render. OpenAI's gpt-oss-safeguard-20b is Apache 2.0, fits in 16 GB and classifies text against a policy you write, which matters because a brand's rules are narrower than a general safety taxonomy. For uploaded reference images, detect faces and route real people to a consent flow; image-to-video on a stranger's photo is the likeliest source of a takedown request. After generation, sample about one frame per second through an image classifier such as Google's ShieldGemma 2, a 4B model covering sexually explicit, dangerous and violent content.
Then label what you ship. C2PA, now at version 2.4, attaches a signed manifest, and its AI and machine learning guidance says generative output should carry the trainedAlgorithmicMedia source type. The open-source c2pa-rs library signs MP4 and MOV files. Re-encoding and screen recording strip metadata, so pair it with an invisible watermark: Meta's VideoSeal is MIT-licensed with a 256-bit model, and the C2PA soft-binding API exists to link the watermark back to the manifest.
The legal floor moved in 2026. Article 50(2) of the EU AI Act applies from 2 August 2026, and the Commission published its Code of Practice on marking and labelling AI-generated content on 10 June 2026. California's AI Transparency Act (Business and Professions Code section 22757 and following, as amended by AB 853) became operative on 2 August 2026 for generative systems with over 1,000,000 monthly users, with a $5,000 civil penalty per violation and each day counted separately. The federal TAKE IT DOWN Act requires covered platforms to remove reported non-consensual intimate imagery, including AI-generated imagery, within 48 hours, enforced from 19 May 2026. None of this is legal advice; all of it belongs in your threat model, next to the rest of LLM security.
- Policy classifier on every prompt, before the queueBlocking a prompt costs almost nothing; blocking a finished clip costs the render.
- Face detection and consent routing on uploaded imagesImage-to-video on real people is the highest-risk path in the product.
- Frame sampling through an image safety classifierAbout one frame per second; a failed clip is deleted, not stored.
- Invisible watermark plus signed C2PA manifest on every filetrainedAlgorithmicMedia source type; the watermark survives the re-encode that strips metadata.
- A 48-hour takedown path with ledger lookupThe TAKE IT DOWN Act clock starts at the request, not at your triage.
- Audio checks on models that generate soundTranscribe and screen speech; skip only if your model is silent.
When should you not self-host AI video generation?
When you render fewer than about 90,000 seconds a month, need audio and top-tier quality today, have nobody to own GPUs, or sell into a market your chosen model's licence excludes. In those cases Veo 3.1 Lite or Fast, Kling 3.0 or Runway Gen-4.5 are usually cheaper, better and less work.
Quality is the uncomfortable part. The strongest models a US or EU company can self-host sit below the best closed ones: with audio, LTX-2.5 scores 1,055 on the Artificial Analysis arena against 1,233 for Gemini Omni Flash and 1,229 for Wan 3.0, both closed. If the feature's value is the wow of the output, a model 178 Elo points behind is a product decision, not an infrastructure one. If its value is volume, consistency and control, that gap matters less than the bill.
There is a middle path most teams should take first. Put one provider interface in front of generation, start on an API, and log prompts, seeds and outcomes from day one. Route previews and drafts to a self-hosted distilled model when volume justifies two GPUs, and keep finals on the API until your own evaluations say otherwise. When the next retirement notice arrives, you change a routing rule instead of a product.
If you do decide to build it, this is the kind of system Axionry builds at $0: the work is split into checkpoints with acceptance criteria agreed before work starts, and each one is invoiced only after you have seen it and accepted it. The serving and privacy side sits under private LLM infrastructure.
Veo 3.1 Lite at $0.03 to $0.05 a second or Veo 3.1 Fast at $0.10 costs less than an $8,718 monthly floor.
Data control is a reason that holds at any volume. Budget the floor and the people.
Apache 2.0, no revenue threshold, and nobody can retire it on you.
Production use above the threshold needs a paid agreement under the LTX-2.x licence.
Both licences exclude those territories, and MiniMax H3 also excludes the US.
Self-hosting AI video generation: common questions
→How much does it cost to self-host an AI video generation model?
Modelled, about $8,718 a month before the first clip: two RTX PRO 6000 GPUs on RunPod at $3,051, four engineer-days of upkeep at $4,000 and an eight-week build amortised at $1,667. On top of that floor, a 4-step distilled Wan 2.2 costs roughly $0.0065 per generated second at 40% utilisation, using RunPod's $2.09 hourly price and 22.5 seconds per five-second clip.
→Which open-source video model is best for commercial use in 2026?
For a US or EU company, LTX-2.5 or Wan 2.2. LTX-2.5 matches Veo 3.1 Lite in the Artificial Analysis silent-video arena, generates audio, and is free below $10 million annual revenue. Wan 2.2 is Apache 2.0 with no thresholds but is silent and about 100 Elo lower. MiniMax H3 ranks higher, but its licence excludes the US, EU, UK and South Korea.
→What GPU do I need to run Wan 2.2 or LTX-2.5?
Stock Wan 2.2 A14B at 720p needs an 80 GB card such as an H100 or A100 80GB. Distilled FP8 builds run on 24 GB cards, and the fastest 4-step NVFP4 build needs a Blackwell GPU such as an RTX 5090 or RTX PRO 6000. Lightricks states a 16 GB minimum for LTX-2.5. Benchmark on the exact card you plan to rent.
→Is self-hosting video generation cheaper than Veo, Kling or Runway?
Only above a volume threshold. With a modelled monthly floor of $8,718, self-hosting beats Veo 3.1 Fast at $0.10 a second above about 87,000 seconds of video a month, and Veo 3.1 Lite at $0.03 above about 290,000. Below that the API is cheaper, and stock Wan 2.2 on one rented H100 costs more per second than Veo 3.1 with audio.
→Do I have to watermark AI-generated video?
In the EU, yes: Article 50(2) of the AI Act has required providers to mark AI-generated video in a machine-readable, detectable way since 2 August 2026. California's AI Transparency Act requires latent disclosures from generative systems with over 1,000,000 monthly users from the same date. A signed C2PA manifest plus an invisible watermark such as VideoSeal covers both approaches.
→What happens if the video API I use is discontinued?
You re-tune every template on another model, on the vendor's timeline. OpenAI announced on 24 March 2026 that the Sora 2 API would shut on 24 September 2026, with no replacement listed. Keep prompts, seeds and reference assets in a provider-neutral format behind one interface, or self-host an Apache 2.0 model such as Wan 2.2 that nobody can retire.
Open the article in your assistant with one click and ask it how this applies to your product.