Gemini 3.6 Flash on AI Gateway, Laguna S 2.1, and agent tools shipping today
TL;DR
Vercel added Gemini 3.6 Flash and Laguna S 2.1 to AI Gateway while new agent platforms and hardware updates rolled out, giving builders faster model testing and SMBs simpler ways to run AI tasks without extra setup.
What shipped
On 21 July new model releases and platform updates centered on faster agent workflows and easier access to coding models. Vercel expanded its AI Gateway with several Gemini and Poolside options while hardware and agent tools appeared across multiple vendors. The changes focus on reducing setup time for testing and running AI in production or daily operations.
Vendor launches
Vercel led with four updates that let teams test and buy AI models through one gateway. Poolside's Laguna S 2.1 and two Gemini Flash variants joined the platform with large context windows aimed at coding and agent tasks. NVIDIA added rack-scale systems for large AI factories while Google released new Gemini models.
- •Searchable on Vercel Searchable uses Vercel's AI SDK and Gateway to ship customer-requested features in 30 minutes while processing over 100 billion tokens, letting brands track visibility in ChatGPT and Perplexity without handling model keys.
- •Laguna S 2.1 on AI Gateway Poolside released Laguna S 2.1 on Vercel AI Gateway in free 256K and paid 1M context versions, giving developers an open-weight model for agentic coding and long-running test tasks.
- •Vercel MCP purchases Vercel MCP now handles Pro plan upgrades, v0 credits, and domain buys directly in chat, letting teams complete payments without leaving the interface.
- •Gemini 3.6 Flash on Gateway Google added Gemini 3.6 Flash and 3.5 Flash-Lite to Vercel AI Gateway, cutting token use on coding and web tasks compared with prior Flash versions.
- •NVIDIA Spectrum-6 NVIDIA launched Spectrum-6 networking for gigascale AI factories that combine hundreds of thousands of GPUs, improving token generation speed in large training clusters.
- •NVIDIA Vera Rubin Vera Rubin NVL72 racks entered production at CoreWeave, Google Cloud, and Azure, delivering lower token cost per watt for partners running frontier models.
- •Gemini 3.6 Flash release Google released Gemini 3.6 Flash and 3.5 Flash-Lite that improve agentic coding output while using fewer model calls than earlier Flash tiers.
- •Gemini 3.5 Flash Cyber Google added a Gemini 3.5 Flash Cyber variant focused on security-related agent tasks alongside the other Flash releases.
Hugging Face trending
Poolside's Laguna-S-2.1 led trending models on Hugging Face as teams downloaded it for text generation and agent work. Other releases included an image-text model, a robotics model, and a low-latency EEG network for edge devices.
- •Laguna-S-2.1 on Hugging Face Poolside's Laguna-S-2.1 text-generation model trended on Hugging Face, offering 1M context for fine-tuning and inference in agentic coding projects.
- •Inkling-NVFP4 on Hugging Face Thinking Machines released Inkling-NVFP4, an image-text-to-text model now available for download and inference through the Hub.
- •MiniCPM-RobotManip on Hugging Face OpenBMB's MiniCPM-RobotManip robotics model trended, supporting fine-tuning for manipulation tasks on standard transformers setups.
- •Motif-3-Beta on Hugging Face Motif Technologies released Motif-3-Beta, a text-generation model trending for fine-tuning and inference on the Hub.
Product Hunt picks
Four new tools launched on Product Hunt that turn web pages into agent instructions, edit video with AI timelines, and handle email or PCB design without heavy coding.
- •OpenChatCut OpenChatCut launched as an open-source AI video editor with a real timeline, letting users cut and edit footage through agent commands.
- •Skim Skim released a free open-source AI email client for Windows that summarizes and sorts messages without manual rules.
- •Manifest Manifest turns any webpage into an action manifest that AI agents can read and execute directly.
- •ProtoFlow ProtoFlow released an AI-powered PCB design tool that helps hardware teams generate board layouts from text prompts.
Other
AWS and Databricks shared new agent and data workflows while OpenRouter added Laguna S 2.1 and Together AI partnered with Y Combinator on a dedicated GPU cluster. Lambda published guidance on choosing coding harnesses that affect model output quality.
- •Amazon Nova self-distillation AWS showed how self-distilled reasoning improves supervised fine-tuning results with Amazon Nova on reasoning benchmarks.
- •Genie coefficient metric IEEE Spectrum proposed the Genie coefficient to measure how well AI follows unspoken user intent beyond standard benchmarks.
- •R&D data in lakehouse Databricks explained why agents need R&D data stored in lakehouses at joint ventures like cellcentric for reliable retrieval.
- •Coding harness evaluation Lambda warned that the same model can produce very different results depending on the harness, urging builders to test multiple options.
- •Laguna S 2.1 on OpenRouter Poolside added Laguna S 2.1 to OpenRouter with 1,049K context at $0.10 per million input tokens for quick testing.
- •Laguna S 2.1 free on OpenRouter Poolside released a free Laguna S 2.1 variant on OpenRouter with 262K context at zero cost for initial agent experiments.
- •Together AI YC cluster Together AI and Y Combinator announced a dedicated GPU cluster for YC startups to run models without shared queue delays.
Industry news
Jack Dorsey launched Buzz as a Slack alternative that includes AI agents in team chats. Other reports covered AI's role in universal entertainment apps, optical chips to extend Moore's Law, and an AI assistant that raised judge productivity in Pakistan.
- •Buzz by Jack Dorsey Jack Dorsey released Buzz, a group chat platform that places human teams and their AI agents in the same conversation thread.
- •Universal entertainment apps AI is pushing Spotify, Netflix, and YouTube toward single apps that handle music, video, podcasts, and recommendations together.
- •Pat Gelsinger optical chips Former Intel CEO Pat Gelsinger is developing light-based chips to increase AI compute power beyond current silicon limits.
- •JudgeGPT in Pakistan A field trial showed JudgeGPT raised case resolution 6.3 percent for Pakistani judges, returning up to $38.50 per dollar when training was provided.
What this means for you
For Vibe Builders: You can now test Laguna S 2.1 and Gemini 3.6 Flash directly through Vercel AI Gateway or OpenRouter without managing keys. Tools like Manifest and OpenChatCut let you turn pages or video into agent actions in minutes. Start with the free Laguna tier on OpenRouter to run a scoped agent task this week.
For Non-techies: New Gemini Flash models and agent chat tools like Buzz reduce the steps needed to get AI to handle email, video, or research tasks. Platforms such as Skim and Manifest remove setup so you can test one workflow today without new accounts. Focus on one concrete job like summarizing messages or turning a page into steps.
For Developers: Vercel AI Gateway now hosts Laguna S 2.1 and Gemini 3.6 Flash with 1M context, letting you swap models in existing code without new SDKs. Evaluate coding harnesses as Lambda recommends before moving any model into production pipelines. Watch the Together AI YC cluster and OpenRouter pricing for cost signals on agent runs.
What to watch next
Watch for production numbers from the new Gemini Flash models on Vercel this week. Track early user reports on Buzz agent chat and the free Laguna tier on OpenRouter. Check Hugging Face for updates to the trending robotics and EEG models.
Harsh’s take
The day showed heavy concentration on model access through gateways rather than new capabilities. Most releases reduce friction for testing but still require builders to pick the right harness or context size to see gains. Hardware updates from NVIDIA target only the largest clusters, leaving smaller teams with the same token-cost tradeoffs.
The practical move is to pick one new model from the Gateway list, run a fixed agent task three times with different context settings, and record token use and output quality before adopting it further.
by Harsh Desai
Sources
Vendor launches
- •How Searchable ships customer-requested features in 30 minutes on Vercel
- •13 Google tips for a fun, productive summer off from college
- •Laguna S 2.1 is now available on AI Gateway
- •Vercel MCP now supports purchases
- •Built for Vera Rubin, NVIDIA Spectrum-6 Arrives in Gigascale AI Factories
- •NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide
- •Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are now available on AI Gateway
- •Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
- •We’re announcing the Alliance for America’s Skilled Trades.
Hugging Face trending
- •Laguna-S-2.1 by poolside trends on HuggingFace
- •Inkling-NVFP4 by thinkingmachines trends on HuggingFace
- •MiniCPM-RobotManip by openbmb trends on HuggingFace
- •Motif-3-Beta by Motif-Technologies trends on HuggingFace
- •Differentiable Logic Gate Networks for Low-Latency EEG Classification on Edge Devices
Product Hunt picks
Other
- •Exploring self-distilled reasoning for supervised fine-tuning with Amazon Nova
- •Why AI Needs a “Genie Coefficient”
- •Why R&D Data Belongs in the Lakehouse - and Why Agents Need It There
- •How Dow Built a Carbon Footprint Ledger on Databricks to Accelerate Sustainability at Scale
- •Your coding harness shouldn't be a black box
- •Poolside: Laguna S 2.1 added on OpenRouter
- •Poolside: Laguna S 2.1 (free) added on OpenRouter
- •🤝 Together AI & Y Combinator announce partnership to deliver the first dedicated YC GPU cluster →
Industry news
- •Jack Dorsey is taking on Slack with Buzz, a group chat platform for teams and their AI agents
- •AI and the rise of the universal entertainment app
- •This Former Intel CEO Wants to Jumpstart Moore’s Law With Light
- •An AI system helped Pakistani judges clear massive backlogs at $38.50 return per dollar invested
More AI news
- Model ReleaseAlibaba's Qwen Audio 3.0 TTS Plus tops Artificial Analysis' Speech Arena
Qwen Audio 3.0 TTS Plus leads the Speech Arena leaderboard with support for 16 languages and natural language style control. It generates speech at 16 characters per second, slower than Sonic 3.5 and Simba 3.2.
- FeatureLovable adds email alerts for sign-ins from new devices and countries
Lovable sends security alert emails for sign-ins from unseen devices and countries, detailing location, browser, OS, and sign-in method.
- FeatureLovable adds auto-deletion for abandoned Enterprise projects
Lovable now lets Enterprise workspaces configure auto-deletion for abandoned projects. Admins set inactivity thresholds while owners receive warnings 5 days and 1 day before deletion.