Stability AI laid off 10% of its staff, which equals approximately 20 employees, following a period of unsustainable growth. This decision follows reports that founder Emad Mostaque left the company after an investor revolt. In 2022, Stability AI secured $101 million in funding, which Bloomberg valued the company at $1 billion. The company also faced scrutiny when Forbes reported that Mostaque made exaggerated statements about his background and the startup. At the same time, some researchers disputed the claim that Stability AI created the Stable Diffusion image generator, noting it was an open-source project developed by various researchers. This downsizing occurs as the broader industry experiences significant shifts. In the first nine months of 2026, more than 170,000 roles were cut globally due to AI. This figure exceeds the 150,000 cuts reported in 2025.
Black Forest Labs growth and market position
Black Forest Labs raised $300 million at a $3.25 billion valuation, making it one of the largest investments in a European AI startup this year. The company, founded in 2024 by Robin Rombach, Andreas Blattmann, and Patrick Esser, has a total of $450 million in funding. The founders, who previously worked at Stability AI, maintain their headquarters in Freiburg, Germany, and a lab in San Francisco. As of August 2026, the startup employs 100 people, with 30 staff members located in San Francisco. The company manages deals with Adobe, Canva, Meta, and Microsoft. Although xAI previously used Black Forest Labs technology for Grok, the startup declined a recent request to license technology again because of xAI’s chaotic work environment. This expansion in Europe follows a trend where the continent saw $5.2 billion invested in AI startups in Q3, an increase from $2 billion in the same quarter last year.
Technical comparison of video models
Animate Diff functions as an extension for Stable Diffusion, adding a motion module to the UNet architecture. This module uses a dataset of video clips to generate consistent motion across frames. Users use Animate Diff to add motion to existing Stable Diffusion checkpoints and LoRAs, which allows for high stylistic flexibility. It produces short clips of 16 to 32 frames but often struggles with motion consistency over long sequences. Generating videos at higher resolutions requires a dedicated GPU with at least 12GB of VRAM. Stable Video Diffusion (SVD) uses end-to-end training on video datasets to generate clips from text or images. SVD-XT produces 25 frames at 576×1024 resolution. While Animate Diff allows for high stylistic variety using the Stable Diffusion ecosystem, SVD focuses on naturalistic and realistic motion. Stability AI also released SV4D 2.0, which produces 48 frames at 576×576 resolution from a 12-frame input video. This model improves spatio-temporal consistency compared to previous 4D versions. You know the difference between a motion module and a dedicated video model.
Market shifts after the Sora shutdown
OpenAI shut down Sora on March 24, 2026, because of unsustainable compute costs. After the sudden shutdown of OpenAI’s Sora in March 2026, the industry saw a rapid migration toward more efficient models like Google Veo 3.1 and Kling AI 3.0 to fill the massive void in high-fidelity video generation. Google Veo 3.1 provides 4K 60fps output and native audio, including dialogue and environmental sounds. Kling AI 3.0 allows for 2-minute videos for $6.99 per month and handles text rendering well. Runway Gen-4.5 targets filmmakers with cinematic quality. Regulatory pressure also hit the sector, as a Dutch court ruled on March 26 that xAI must stop creating sexualized images through Grok, with fines reaching €100,000 daily. This ruling follows data from the UK’s Internet Watch Foundation showing an increase in AI-generated child sexual abuse imagery in 2025. The UK reported 8,029 such images, a 14% increase from the previous year.
Physics and character consistency in 2026
MiniMax’s Hailuo AI uses a physics-optimized rendering pipeline that cuts cloud computing expenses by 52% while maintaining 4K output quality. This July 2026 update reduces rendering artifacts by 63%. Digen AI Agent uses autonomous material property adjustments, which reduces manual editing time by 78%. Researchers in 2026 found that hybrid architectures combining traditional physics engines with neural networks achieve 89% simulation accuracy at 22% of the computational cost of pure physics-based approaches. NVIDIA’s H200 Tensor Core GPUs deliver 2.3x better performance per watt than 2025 hardware, which lowers operational costs for providers. Character consistency also improved, with Google’s $12/month tier reducing identity drift by 67% compared to 2025 models. SoulGen 2.0 achieved a 0.96 ID consistency score for facial recognition across video sequences.
Subscription pricing and European investment
The average price for one minute of 1080p video fell from $4.20 in 2025 to $2.75 in the third quarter of 2026. Google’s May 2026 I/O announcement included a 27% price reduction for enterprise models, and its consumer tiers start at $8.99 for 720p output. MiniMax offers professional animation at $0.12 per generated minute. Most affordable plans include at least 3 character slots with 89% consistency ratings. In Europe, Mistral AI announced a $2 billion Series C round in early September, while Nscale raised $1.1 billion in September 2025. Digen AI Agent now handles 83% of editing workflows automatically.
| Feature | Animate Diff | Stable Video (SVD) | FLUX.3 |
|---|---|---|---|
| Primary Type | Motion Module | Dedicated Model | Text-to-Video |
| Mechanism | UNet integration | End-to-end training | Diffusion |
| Max Resolution | Depends on Base | 576×1024 (SVD-XT) | 4K (via specialized modes) |
| Best Use Case | Stylized Animation | Realistic Motion | Robotics and Video-Action |
Physical AI and the expansion of FLUX.3
Black Forest Labs released FLUX.3 in July 2026, which includes the FLUX3 x mimic model. This video-action model runs on robots and performs soft-body manipulation work at Audi car production facilities. The company also works with hardware firms on features for smart glasses and robots. The company’s latent diffusion approach uses fewer resources than competitors. This technology allows for powerful models that require orders of magnitude less compute. Black Forest Labs continues to compete with better funded Silicon Valley companies by maintaining a research-focused approach. The company also maintains a presence in the image editing market through partnerships with Adobe and Canva.
Will the reliance on latent diffusion models eventually encounter the same physical inconsistencies that plagued earlier generative models?
