Gemini 3.8 Flash lands on AI Gateway at 50% off, Muse Spark 1.3 and GLM-5.3 deals, plus agent tools to test
TL;DR
Vercel AI Gateway added three major models with discounts and improved agent performance while new coding agents and video tools appeared on Product Hunt and Fal.
What shipped
On 2 September multiple vendors released updated models and agent features aimed at faster coding and multi-step tasks. The updates center on lower-cost access through existing gateways and new practical tools for builders. Industry coverage also examined detection challenges and safety concerns around reasoning techniques.
Vendor launches
Vercel AI Gateway added GLM-5.3 at half price through DigitalOcean, Muse Spark 1.3 from Meta, and Gemini 3.8 Flash from Google with a year-end discount. These releases focus on agent work, coding, and longer context windows. Other Google announcements covered partnerships and internal programs but stayed secondary to the model updates.
- •GLM-5.3 discount GLM-5.3 runs at 50% off on AI Gateway via DigitalOcean until 8 September using the promo code zai/glm-5.3-promo-50, after which standard routing across providers resumes.
- •Muse Spark 1.3 release Meta released Muse Spark 1.3 on AI Gateway with a 1M token window and better coding efficiency that reduces turns and filler text compared with the prior version.
- •Gemini 3.8 Flash on Gateway Google added Gemini 3.8 Flash to AI Gateway at 50% off through December with 1M context, tool calling, web search, and default thinking for software engineering tasks.
- •Gemini 3.8 Flash Cyber Google introduced Gemini 3.8 Flash Cyber for agentic workflows and cybersecurity use cases alongside the standard Flash model.
Hugging Face trending
A 4B text-generation model from XHToken appeared in the trending list on Hugging Face Hub. A research paper on knowledge distillation during mid-training also gained attention for its findings on reasoning versus recall.
- •Spark-X2.5-4B model XHToken released Spark-X2.5-4B, a 4B text-generation model now trending on Hugging Face and ready for download, fine-tuning, and inference.
- •Knowledge distillation study New experiments show logit-based distillation during mid-training improves reasoning more than factual recall in smaller language models.
Fal model gallery
H3 Max reference-to-video: Fal released H3 Max, a MiniMax H3 variant optimized for stylized reference-to-video generation with improved throughput and prompt following.
Product Hunt picks
Four new AI tools launched on Product Hunt focused on embroidery digitizing, local Grok control, advanced Claude coding, and collaborative AI design canvases.
- •Stitch AI Dynamic Mockups launched Stitch AI, an embroidery digitizing agent that converts designs into stitch files.
- •Porte Porte lets users control local Grok sessions directly from a phone interface.
- •Claude Fable 5.1 Anthropic released Claude Fable 5.1 tuned for coding and knowledge work.
- •Doop Doop provides a shared canvas where multiple AI agents collaborate on design tasks in real time.
Industry news
Coverage examined AI detection limits, OpenAI training policy, Gemini 3.8 Flash benchmark details, and new reasoning techniques. A Russian startup demonstrated wordless model-to-model communication.
- •AI detection challenges Pangram CEO noted that distinguishing AI text in job applications and reviews is harder than simple real-or-fake checks.
- •US OpenAI brief The US government filed a brief supporting OpenAI in the ongoing debate over training LLMs on copyrighted material.
- •Gemini 3.8 Flash analysis The Decoder reported Gemini 3.8 Flash matches some Claude Opus 5 agentic coding scores yet uses 30% more output tokens due to its reasoning style.
- •llm-gemini 0.34 update Simon Willison released llm-gemini 0.34 adding Gemini 3.8 Flash support with adjustable thinking levels.
- •Mostik model communication Russian startup Mostik demonstrated AI models exchanging information without using words.
- •OpenAI recurrent depth OpenAI's Astra model uses recurrent depth reasoning that operates outside standard sequential steps, raising safety questions.
- •Dead internet concerns Pangram warned that AI-generated content in reviews and claims is pushing the internet closer to dead-internet conditions.
Other
Databricks expanded Genie Agents, GitHub shared Copilot cost tips, AWS posted Bedrock examples, and OpenRouter added Muse Spark 1.3 at two price tiers.
- •Genie Agents expansion Databricks extended Genie Spaces into Genie Agents for deeper file reasoning and analysis at the Data and AI Summit.
- •GitHub Copilot efficiency GitHub published methods to cut wasted output tokens in Copilot without lowering task completion rates.
- •AWS Bedrock dashboard checks An AWS team described using Amazon Bedrock to detect dashboard content failures at scale.
- •Cohere small models post Cohere outlined how smaller models deliver enterprise impact at lower cost than large frontier systems.
- •Muse Spark 1.3 on OpenRouter Meta added Muse Spark 1.3 to OpenRouter with 1,049k context at $1.25 per million input tokens.
- •Muse Spark 1.3 Contributor tier OpenRouter now lists the lower-cost Muse Spark 1.3 Contributor tier at $0.10 input and $0.20 output per million tokens.
- •AWS OpenAI cross-region AWS enabled global cross-Region inference for OpenAI models on Bedrock from Australia.
Replicate new models
ltxvideo-2.3-lora: Replicate released ltxvideo-2.3-lora for community LoRA tasks including ingredients and camera control via its HTTP API.
What this means for you
For Vibe Builders: You can now route agent and coding work through Vercel AI Gateway using Gemini 3.8 Flash or Muse Spark 1.3 at discounted rates and test the new Doop canvas or Stitch AI digitizer without writing code. The 1M context windows and tool calling reduce the number of steps needed for multi-turn tasks. Add the model IDs to an existing gateway setup and compare output length against your current stack this week.
For Non-techies: Discounted access to Gemini 3.8 Flash and Muse Spark 1.3 on AI Gateway means lower costs for everyday document and image tasks that previously required multiple prompts. Product Hunt tools like Porte for phone control of local models and Doop for shared AI design canvases let small teams start without new accounts. Watch for the Muse Spark Contributor tier if you need the cheapest entry point.
For Developers: Gemini 3.8 Flash, Muse Spark 1.3, and GLM-5.3 now sit behind one gateway with documented token pricing and thinking controls. Evaluate the 30% higher output token use on Gemini 3.8 Flash against your existing agent benchmarks before switching production traffic. Track the OpenRouter Contributor tier and Replicate LoRA endpoints for cheap local testing before committing to any single provider.
What to watch next
Watch for production benchmarks on Gemini 3.8 Flash token usage this week and any follow-up releases from OpenAI on recurrent depth. Check whether the Muse Spark Contributor tier stays at the listed OpenRouter price after the initial launch window.
Harsh’s take
The day showed repeated budget model drops from Google and Meta while frontier releases stayed absent. Gateway discounts and contributor tiers lower the barrier for testing yet also highlight that raw speed claims often hide higher real-world token spend. Builders should pick one new model from the Gateway list, run a fixed agent task set, and measure total tokens and latency against their current default before the discounts end.
by Harsh Desai
Sources
Vendor launches
- •GLM-5.3 is 50% off through DigitalOcean on AI Gateway
- •Muse Spark 1.3 now available on AI Gateway
- •Free domain with Pro offer now includes .app and .dev
- •Build a measurement stack you can rely on to steer your campaigns.
- •Gemini 3.8 Flash now available on AI Gateway
- •MrBeast partners with Gemini to turn impossibly big ideas into reality
- •Introducing Gemini 3.8 Flash and 3.8 Flash Cyber
- •Proactive cyber defense for governments and enterprises
- •Book insights in Google Play Books is now available in more than a million ebooks and in the iOS app.
- •Our new partnership brings long-duration energy storage to West Virginia.
Hugging Face trending
- •Spark-X2.5-4B by XHToken trends on HuggingFace
- •Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
Fal model gallery
Product Hunt picks
Industry news
- •Pangram’s Max Spero on why AI detection is harder than ‘Real or Fake’
- •We’re ‘dangerously close’ to dead internet theory, says Pangram’s CEO
- •US government sides with OpenAI on issue of training LLMs on copyrighted material
- •Gemini 3.8 Flash is Google's third budget model in six weeks while frontier models remain MIA
- •llm-gemini 0.34
- •These Russian Mathematicians Taught AI Models How to Talk to Each Other Without Using Words
- •OpenAI’s new reasoning technique alarms AI safety experts
Other
- •Expanding Genie Agents: Deep analysis, file reasoning, and more
- •How we make AI coding more cost efficient without sacrificing task quality
- •How an AWS team detects dashboard content failures at scale using Amazon Bedrock
- •How small AI models can make a big impact for enterprisesSep 02, 20,267 min read
- •Skip the reorg: Megan Filbin on where value lives when AI does the work
- •Meta: Muse Spark 1.3 now available on OpenRouter (1,049k context, $1.25/M in, $4.25/M out)
- •Meta: Muse Spark 1.3 Contributor added on OpenRouter
- •Accessing OpenAI models on Amazon Bedrock from Australia with global cross-Region inference
- •Decoding the new AI lingo: Loops, harnesses, squads, hill climbing… oh my!
Replicate new models
More AI news
- Daily RoundupMercury 2.5 and Nex-N2.5 on OpenRouter, Meta Muse agent, and Vercel speed gains for agents
Model releases and agent tools expand options for builders while enterprise deals and routing improvements reduce friction for teams shipping AI products today.
- Daily RoundupHugging Face model wave, Fal H3 Turbo video, and Product Hunt AI agents
Hugging Face saw five models trend including text, image, video and speech tools while Fal released an upgraded text-to-video model and Product Hunt featured new agent-style apps for note-taking and local coding.
- Weekly DigestHermes Agent Bot Mode and OpenClaw 2.0 add group chats plus desktop controls
Hermes Agent and OpenClaw both shipped major updates this week that turn single agents into coordinated groups and give them direct control over browsers, desktops, and team workflows.