Skip to content
Gemini 3.8 Flash lands on AI Gateway at 50% off, Muse Spark 1.3 and GLM-5.3 deals, plus agent tools to test | My AI Guide

Gemini 3.8 Flash lands on AI Gateway at 50% off, Muse Spark 1.3 and GLM-5.3 deals, plus agent tools to test

By Harsh Desai
Share

TL;DR

Vercel AI Gateway added three major models with discounts and improved agent performance while new coding agents and video tools appeared on Product Hunt and Fal.

What shipped

On 2 September multiple vendors released updated models and agent features aimed at faster coding and multi-step tasks. The updates center on lower-cost access through existing gateways and new practical tools for builders. Industry coverage also examined detection challenges and safety concerns around reasoning techniques.

Vendor launches

Vercel AI Gateway added GLM-5.3 at half price through DigitalOcean, Muse Spark 1.3 from Meta, and Gemini 3.8 Flash from Google with a year-end discount. These releases focus on agent work, coding, and longer context windows. Other Google announcements covered partnerships and internal programs but stayed secondary to the model updates.

  • GLM-5.3 discount GLM-5.3 runs at 50% off on AI Gateway via DigitalOcean until 8 September using the promo code zai/glm-5.3-promo-50, after which standard routing across providers resumes.
  • Muse Spark 1.3 release Meta released Muse Spark 1.3 on AI Gateway with a 1M token window and better coding efficiency that reduces turns and filler text compared with the prior version.
  • Gemini 3.8 Flash on Gateway Google added Gemini 3.8 Flash to AI Gateway at 50% off through December with 1M context, tool calling, web search, and default thinking for software engineering tasks.
  • Gemini 3.8 Flash Cyber Google introduced Gemini 3.8 Flash Cyber for agentic workflows and cybersecurity use cases alongside the standard Flash model.

Hugging Face trending

A 4B text-generation model from XHToken appeared in the trending list on Hugging Face Hub. A research paper on knowledge distillation during mid-training also gained attention for its findings on reasoning versus recall.

  • Spark-X2.5-4B model XHToken released Spark-X2.5-4B, a 4B text-generation model now trending on Hugging Face and ready for download, fine-tuning, and inference.
  • Knowledge distillation study New experiments show logit-based distillation during mid-training improves reasoning more than factual recall in smaller language models.

Fal model gallery

H3 Max reference-to-video: Fal released H3 Max, a MiniMax H3 variant optimized for stylized reference-to-video generation with improved throughput and prompt following.

Product Hunt picks

Four new AI tools launched on Product Hunt focused on embroidery digitizing, local Grok control, advanced Claude coding, and collaborative AI design canvases.

  • Stitch AI Dynamic Mockups launched Stitch AI, an embroidery digitizing agent that converts designs into stitch files.
  • Porte Porte lets users control local Grok sessions directly from a phone interface.
  • Claude Fable 5.1 Anthropic released Claude Fable 5.1 tuned for coding and knowledge work.
  • Doop Doop provides a shared canvas where multiple AI agents collaborate on design tasks in real time.

Industry news

Coverage examined AI detection limits, OpenAI training policy, Gemini 3.8 Flash benchmark details, and new reasoning techniques. A Russian startup demonstrated wordless model-to-model communication.

  • AI detection challenges Pangram CEO noted that distinguishing AI text in job applications and reviews is harder than simple real-or-fake checks.
  • US OpenAI brief The US government filed a brief supporting OpenAI in the ongoing debate over training LLMs on copyrighted material.
  • Gemini 3.8 Flash analysis The Decoder reported Gemini 3.8 Flash matches some Claude Opus 5 agentic coding scores yet uses 30% more output tokens due to its reasoning style.
  • llm-gemini 0.34 update Simon Willison released llm-gemini 0.34 adding Gemini 3.8 Flash support with adjustable thinking levels.
  • Mostik model communication Russian startup Mostik demonstrated AI models exchanging information without using words.
  • OpenAI recurrent depth OpenAI's Astra model uses recurrent depth reasoning that operates outside standard sequential steps, raising safety questions.
  • Dead internet concerns Pangram warned that AI-generated content in reviews and claims is pushing the internet closer to dead-internet conditions.

Other

Databricks expanded Genie Agents, GitHub shared Copilot cost tips, AWS posted Bedrock examples, and OpenRouter added Muse Spark 1.3 at two price tiers.

  • Genie Agents expansion Databricks extended Genie Spaces into Genie Agents for deeper file reasoning and analysis at the Data and AI Summit.
  • GitHub Copilot efficiency GitHub published methods to cut wasted output tokens in Copilot without lowering task completion rates.
  • AWS Bedrock dashboard checks An AWS team described using Amazon Bedrock to detect dashboard content failures at scale.
  • Cohere small models post Cohere outlined how smaller models deliver enterprise impact at lower cost than large frontier systems.
  • Muse Spark 1.3 on OpenRouter Meta added Muse Spark 1.3 to OpenRouter with 1,049k context at $1.25 per million input tokens.
  • Muse Spark 1.3 Contributor tier OpenRouter now lists the lower-cost Muse Spark 1.3 Contributor tier at $0.10 input and $0.20 output per million tokens.
  • AWS OpenAI cross-region AWS enabled global cross-Region inference for OpenAI models on Bedrock from Australia.

Replicate new models

ltxvideo-2.3-lora: Replicate released ltxvideo-2.3-lora for community LoRA tasks including ingredients and camera control via its HTTP API.

What this means for you

For Vibe Builders: You can now route agent and coding work through Vercel AI Gateway using Gemini 3.8 Flash or Muse Spark 1.3 at discounted rates and test the new Doop canvas or Stitch AI digitizer without writing code. The 1M context windows and tool calling reduce the number of steps needed for multi-turn tasks. Add the model IDs to an existing gateway setup and compare output length against your current stack this week.

For Non-techies: Discounted access to Gemini 3.8 Flash and Muse Spark 1.3 on AI Gateway means lower costs for everyday document and image tasks that previously required multiple prompts. Product Hunt tools like Porte for phone control of local models and Doop for shared AI design canvases let small teams start without new accounts. Watch for the Muse Spark Contributor tier if you need the cheapest entry point.

For Developers: Gemini 3.8 Flash, Muse Spark 1.3, and GLM-5.3 now sit behind one gateway with documented token pricing and thinking controls. Evaluate the 30% higher output token use on Gemini 3.8 Flash against your existing agent benchmarks before switching production traffic. Track the OpenRouter Contributor tier and Replicate LoRA endpoints for cheap local testing before committing to any single provider.

What to watch next

Watch for production benchmarks on Gemini 3.8 Flash token usage this week and any follow-up releases from OpenAI on recurrent depth. Check whether the Muse Spark Contributor tier stays at the listed OpenRouter price after the initial launch window.

Harshs take

The day showed repeated budget model drops from Google and Meta while frontier releases stayed absent. Gateway discounts and contributor tiers lower the barrier for testing yet also highlight that raw speed claims often hide higher real-world token spend. Builders should pick one new model from the Gateway list, run a fixed agent task set, and measure total tokens and latency against their current default before the discounts end.

by Harsh Desai

Sources

Vendor launches

Hugging Face trending

Fal model gallery

Product Hunt picks

Industry news

Other

Replicate new models

More AI news

Everything AI. One email.
Every Monday.

New tools. Model launches. Plugins. Repos. Tactics. The moves the sharpest builders are making right now, before everyone else.

No spam. Unsubscribe anytime.