Alibaba has opened the beta of Wan 3.0, a video model that doubles the maximum clip length and swallows entire documents as input. Thirty seconds at 1080p costs six dollars, and a PDF now works as a storyboard the same way a written prompt does.
Key Takeaways
- Wan 3.0 generates clips of up to thirty seconds, against fifteen for Wan 2.5.
- The model reads text, images, video, audio, PDFs, web pages and slide decks in a single call.
- The beta runs on the Wan platform, on Alibaba Cloud Model Studio and through the Qwen Cloud API.
Have an AI Sum Up This Article
ChatGPTThirty Seconds and Inputs Nobody Expected
Alibaba pushed Wan 3.0 into public beta, and the headline number is duration. The model returns clips of thirty seconds where Wan 2.5 capped out at fifteen. Doubling sounds incremental, yet it redraws what a video generator can be asked to do, since a fifteen second ceiling rules out most ad formats outright.
The second change sits on the input side. Wan 3.0 takes text, up to ten images, five videos and five audio clips, and it also reads PDFs, web pages and presentation files. Every one of those formats can land in the same generation call, with no conversion step in between.
A sales deck therefore becomes a direct starting point. Instead of writing a prompt that describes a scene, the user hands over a document and lets the model pull the narrative structure out of it. The beta is live on the official platform where Alibaba ships its video models, on Alibaba Cloud Model Studio and through the Qwen Cloud API.
Two quieter features round out the release. The model recommends the clip length that suits the material it was given, and it can extend an existing video rather than regenerate the whole thing. On longer formats that compute saving is anything but theoretical, because every regenerated second is billed.
Visual stability closes the list. Alibaba names face distortion and interface inconsistency as the two targets, and those are the exact defects that give a generated clip away, the moment a logo or a screen shows up inside the frame.
Google opened this multimodal lane earlier in the year with a model able to start from almost any kind of input, though it stopped short of office files. Alibaba moves the line further by treating a working document as one more media type.
What Six Dollars a Clip Changes for Studios
Pricing is published per second, which is still uncommon in this segment. The entry tier starts at five cents a second in 480p Standard and climbs to twenty eight cents in 1080p Prime. Thirty seconds in 1080p Standard lands at six dollars.
That number changes the calculation for teams shipping short formats at volume. Running ten variants of the same clip costs less than an hour of editing time, and the real constraint moves from production budget to selection time.
Document inputs sharpen the effect inside companies. A training team that already keeps its material in slide decks skips the creative brief entirely. It uploads the file, gets a first pass back, and the human work shifts to correction and choice.
The catch is the beta label. Alibaba commits to no availability guarantee, and the published rates apply to the opening phase only. The group ran the same playbook when it shipped its largest language model this summer, opening wide first and tightening later.
For independents, the account is the wall. Access runs through Alibaba Cloud or Qwen Cloud, two environments built for developers, with the identity checks and usage billing that come with them. A solo creator who just wants to try thirty seconds is plainly not who this opening is aimed at.
What nobody can price yet is how good those thirty seconds actually look. A doubled duration exposes coherence failures for twice as long, and that is where early hands-on reports will settle the verdict, well ahead of any benchmark table.
More articles on Horizon
- V4-Flash-Vision Test: DeepSeek Wins on Agent Work
- Ox Alpha: A Free AI Model Arrives With No Vendor
- DeepSeek V4-Flash-Vision Closes In on Opus 4.8
Google, ByteDance and the Race on Duration
Duration has turned into the most readable comparison point between video models. Every step up unlocks a format the previous ceiling forbade, and thirty seconds now covers a full ad, an opening sequence or a short tutorial.
ByteDance has held this ground since the start of the year with a video model that rattled Hollywood. The fight plays out on face rendering and shot consistency, which are precisely the two areas Alibaba highlights for Wan 3.0.
On the American side, the pressure lands on price. A per second rate published this plainly makes comparable what subscription products keep vague, since the real cost of a clip there depends on a credit system whose conversion rate shifts with every update.
Document input is the piece Alibaba pushes hardest. Reading a deck or a web page in the same call as a reference image moves the model out of the creative studio and into internal production, a market where the competitor is no longer a rival generator but an agency.
The editing features point the same way. Reselecting one interval and regenerating only that stretch brings the model closer to an editing timeline, where video generation has so far stayed an all or nothing exercise.
Volume is the open question. A thirty second generation burns far more compute than a fifteen second one, and this beta doubles as a measurement of how many users actually reach for the long format before the pricing grid hardens.
Follow the story on Horizon.


