Most businesses are sitting on hours of useful information that nobody has time to properly process.
Customer calls contain objections and recurring questions. Meetings contain decisions that never make it into the CRM. Support recordings reveal product problems. Training videos hold knowledge that employees struggle to find later. Marketing teams produce video that could become articles, social posts, FAQs and sales material.
Alibaba’s new Qwen3.8-Omni-Flash makes processing all of that information considerably more practical.
The model can understand text, images, audio and video inside the same request, work with up to 1 million tokens of context, use tools and web search, and understand audio across 113 languages and dialects. Alibaba also says the API cost per hour of audio input has fallen by more than 98% compared with its previous Qwen3.5-Omni-Plus model, while audio-video input costs have fallen by more than 93%.
Those numbers are Alibaba’s own comparison with its previous model, so I wouldn’t translate them into “AI audio is now 98% cheaper” generally. The important point is that serious audio and video analysis is moving into a price range where smaller businesses can start using it continuously instead of occasionally.
For SMBs, that is where this release gets interesting.
Your calls and meetings can become business data
Consider a small consultancy with ten client calls every week.
Each call might contain new requirements, objections, deadlines, pricing questions and follow-up actions. Someone usually writes a few notes, creates a task or two and then moves on.
A multimodal model can process the recording and extract much more structure.
It could identify the customer, summarize what was discussed, list decisions, extract promised actions, identify unresolved questions and prepare a CRM update. Because Qwen3.8-Omni-Flash supports function calling, that analysis can become part of an automation that sends the information somewhere useful.
The workflow could look like this:
Call recording → Qwen analysis → structured data → CRM update → Todoist/Jira tasks → follow-up email draft
That is more useful than simply producing a transcript.
The transcript is raw material. The business value comes from turning what happened during the conversation into actions that actually enter the systems running the company.
A solo consultant could use the same workflow on every discovery call without employing someone to document meetings. A small sales team could automatically track recurring objections. A service company could look for customer questions that repeatedly slow down sales.
Support conversations become much easier to analyze
Customer support is another strong use case.
A company handling hundreds of calls, voice notes or recorded support sessions each month has a large dataset explaining what customers struggle with. Smaller businesses rarely have someone available to systematically review that information.
Qwen can process long audio and identify repeated problems.
For example, a SaaS company could analyze a month of customer calls and ask which features create the most confusion, which questions occur most frequently, which issues tend to trigger cancellations and which problems support agents struggle to explain.
A holiday rental company could analyze guest calls and identify recurring questions around check-in, parking, Wi-Fi or air conditioning.
An ecommerce company could examine customer-service recordings and discover that customers repeatedly misunderstand the same delivery policy.
Those findings can then feed back into the business. FAQs can be improved. Website copy can be rewritten. Onboarding can change. Support scripts can be updated.
This is where I think the cheaper processing really matters. An occasional AI analysis of ten calls is useful. An automated system processing every relevant conversation creates a much richer picture of what customers are actually experiencing.
Video becomes usable input for automation
The video side may be just as important for businesses already creating content.
Imagine recording a 20-minute product demonstration.
A multimodal model can understand what is being said and what appears on screen. That makes workflows possible where the video becomes the source material for several other outputs.
A business could process the recording and generate a detailed article brief, extract the products or features demonstrated, identify useful short-form sections, prepare chapter descriptions, create social-post ideas and generate an FAQ based on what was explained.
For a solo entrepreneur creating YouTube content, training material or product demonstrations, one recording can become the raw material for an entire content pipeline.
There is also an internal use case. Companies accumulate onboarding videos, recorded webinars, Zoom meetings and training sessions. Processing that material with a model capable of understanding both audio and visuals makes it easier to extract procedures, create documentation and build searchable internal knowledge.
Qwen’s 1 million-token context window helps with longer and more complex source material, although real production limits will still depend on the exact media format and API configuration.
Multi-speaker audio is particularly useful
Qwen3.8-Omni-Flash also improves multi-speaker recognition and supports multi-channel audio.
That matters because real business recordings are messy.
People interrupt each other. Several people attend a meeting. Customers speak with different accents. A recording may contain both a salesperson and a prospect.
Correctly understanding who said what is essential if the model is expected to turn conversations into tasks or CRM records.
A summary saying “customer will send the contract” when it was actually your salesperson who promised to send it is worse than having no automation at all.
Alibaba reports substantial improvements in multi-speaker and long-form audio understanding compared with its previous generation. I would still test this aggressively on your own recordings before allowing the output to update important systems automatically.
Where I would use it first
For an SMB or solo entrepreneur, I wouldn’t begin by designing a giant multimodal AI platform.
Pick a workflow where audio or video already exists.
A good first experiment would be sales calls. Process 20 real recordings and extract the customer requirements, objections, actions and follow-up information. Compare the results with your own notes.
Another good option is meetings. Generate structured decisions and tasks rather than basic summaries.
Content-heavy businesses could test one video-to-content pipeline. Support-heavy businesses could analyze a month of calls and look for recurring topics.
Measure the result in hours saved and useful information recovered.
The model is available through Alibaba Cloud Model Studio, including a Frankfurt region. European businesses should still review the relevant data-processing, retention and contractual terms before sending customer recordings or sensitive business information into any external AI API.
Why this matters
Text AI became useful to SMBs very quickly because businesses already produce enormous amounts of text.
The same thing is starting to happen with audio and video.
Calls, meetings, webinars, demos, support recordings and videos are full of business information. Historically, extracting that information at scale required transcription services, specialist software and several processing steps.
Models like Qwen3.8-Omni-Flash can understand those formats together and connect the result to tools.
The falling processing cost makes the biggest difference. When analyzing every call becomes affordable, businesses can design workflows around continuous understanding rather than occasional manual review.
For a small company, that means conversations and recordings can become structured, searchable and actionable business data.
My recommendation: Test Qwen3.8-Omni-Flash on one real audio or video workflow where you already produce enough material to make manual analysis unrealistic. Sales calls, customer support, meetings and content repurposing are the strongest starting points. Benchmark its accuracy and actual API cost against your current process before expanding it.
