Skip to content
Gemini 3.6 Flash on AI Gateway, Laguna S 2.1, and agent tools shipping today | Daily AI roundup cover

Gemini 3.6 Flash on AI Gateway, Laguna S 2.1, and agent tools shipping today

By Harsh Desai
Share

TL;DR

Vercel added Gemini 3.6 Flash and Laguna S 2.1 to AI Gateway while new agent platforms and hardware updates rolled out, giving builders faster model testing and SMBs simpler ways to run AI tasks without extra setup.

What shipped

On 21 July new model releases and platform updates centered on faster agent workflows and easier access to coding models. Vercel expanded its AI Gateway with several Gemini and Poolside options while hardware and agent tools appeared across multiple vendors. The changes focus on reducing setup time for testing and running AI in production or daily operations.

Vendor launches

Vercel led with four updates that let teams test and buy AI models through one gateway. Poolside's Laguna S 2.1 and two Gemini Flash variants joined the platform with large context windows aimed at coding and agent tasks. NVIDIA added rack-scale systems for large AI factories while Google released new Gemini models.

  • Searchable on Vercel Searchable uses Vercel's AI SDK and Gateway to ship customer-requested features in 30 minutes while processing over 100 billion tokens, letting brands track visibility in ChatGPT and Perplexity without handling model keys.
  • Laguna S 2.1 on AI Gateway Poolside released Laguna S 2.1 on Vercel AI Gateway in free 256K and paid 1M context versions, giving developers an open-weight model for agentic coding and long-running test tasks.
  • Vercel MCP purchases Vercel MCP now handles Pro plan upgrades, v0 credits, and domain buys directly in chat, letting teams complete payments without leaving the interface.
  • Gemini 3.6 Flash on Gateway Google added Gemini 3.6 Flash and 3.5 Flash-Lite to Vercel AI Gateway, cutting token use on coding and web tasks compared with prior Flash versions.
  • NVIDIA Spectrum-6 NVIDIA launched Spectrum-6 networking for gigascale AI factories that combine hundreds of thousands of GPUs, improving token generation speed in large training clusters.
  • NVIDIA Vera Rubin Vera Rubin NVL72 racks entered production at CoreWeave, Google Cloud, and Azure, delivering lower token cost per watt for partners running frontier models.
  • Gemini 3.6 Flash release Google released Gemini 3.6 Flash and 3.5 Flash-Lite that improve agentic coding output while using fewer model calls than earlier Flash tiers.
  • Gemini 3.5 Flash Cyber Google added a Gemini 3.5 Flash Cyber variant focused on security-related agent tasks alongside the other Flash releases.

Hugging Face trending

Poolside's Laguna-S-2.1 led trending models on Hugging Face as teams downloaded it for text generation and agent work. Other releases included an image-text model, a robotics model, and a low-latency EEG network for edge devices.

  • Laguna-S-2.1 on Hugging Face Poolside's Laguna-S-2.1 text-generation model trended on Hugging Face, offering 1M context for fine-tuning and inference in agentic coding projects.
  • Inkling-NVFP4 on Hugging Face Thinking Machines released Inkling-NVFP4, an image-text-to-text model now available for download and inference through the Hub.
  • MiniCPM-RobotManip on Hugging Face OpenBMB's MiniCPM-RobotManip robotics model trended, supporting fine-tuning for manipulation tasks on standard transformers setups.
  • Motif-3-Beta on Hugging Face Motif Technologies released Motif-3-Beta, a text-generation model trending for fine-tuning and inference on the Hub.

Product Hunt picks

Four new tools launched on Product Hunt that turn web pages into agent instructions, edit video with AI timelines, and handle email or PCB design without heavy coding.

  • OpenChatCut OpenChatCut launched as an open-source AI video editor with a real timeline, letting users cut and edit footage through agent commands.
  • Skim Skim released a free open-source AI email client for Windows that summarizes and sorts messages without manual rules.
  • Manifest Manifest turns any webpage into an action manifest that AI agents can read and execute directly.
  • ProtoFlow ProtoFlow released an AI-powered PCB design tool that helps hardware teams generate board layouts from text prompts.

Other

AWS and Databricks shared new agent and data workflows while OpenRouter added Laguna S 2.1 and Together AI partnered with Y Combinator on a dedicated GPU cluster. Lambda published guidance on choosing coding harnesses that affect model output quality.

  • Amazon Nova self-distillation AWS showed how self-distilled reasoning improves supervised fine-tuning results with Amazon Nova on reasoning benchmarks.
  • Genie coefficient metric IEEE Spectrum proposed the Genie coefficient to measure how well AI follows unspoken user intent beyond standard benchmarks.
  • R&D data in lakehouse Databricks explained why agents need R&D data stored in lakehouses at joint ventures like cellcentric for reliable retrieval.
  • Coding harness evaluation Lambda warned that the same model can produce very different results depending on the harness, urging builders to test multiple options.
  • Laguna S 2.1 on OpenRouter Poolside added Laguna S 2.1 to OpenRouter with 1,049K context at $0.10 per million input tokens for quick testing.
  • Laguna S 2.1 free on OpenRouter Poolside released a free Laguna S 2.1 variant on OpenRouter with 262K context at zero cost for initial agent experiments.
  • Together AI YC cluster Together AI and Y Combinator announced a dedicated GPU cluster for YC startups to run models without shared queue delays.

Industry news

Jack Dorsey launched Buzz as a Slack alternative that includes AI agents in team chats. Other reports covered AI's role in universal entertainment apps, optical chips to extend Moore's Law, and an AI assistant that raised judge productivity in Pakistan.

  • Buzz by Jack Dorsey Jack Dorsey released Buzz, a group chat platform that places human teams and their AI agents in the same conversation thread.
  • Universal entertainment apps AI is pushing Spotify, Netflix, and YouTube toward single apps that handle music, video, podcasts, and recommendations together.
  • Pat Gelsinger optical chips Former Intel CEO Pat Gelsinger is developing light-based chips to increase AI compute power beyond current silicon limits.
  • JudgeGPT in Pakistan A field trial showed JudgeGPT raised case resolution 6.3 percent for Pakistani judges, returning up to $38.50 per dollar when training was provided.

What this means for you

For Vibe Builders: You can now test Laguna S 2.1 and Gemini 3.6 Flash directly through Vercel AI Gateway or OpenRouter without managing keys. Tools like Manifest and OpenChatCut let you turn pages or video into agent actions in minutes. Start with the free Laguna tier on OpenRouter to run a scoped agent task this week.

For Non-techies: New Gemini Flash models and agent chat tools like Buzz reduce the steps needed to get AI to handle email, video, or research tasks. Platforms such as Skim and Manifest remove setup so you can test one workflow today without new accounts. Focus on one concrete job like summarizing messages or turning a page into steps.

For Developers: Vercel AI Gateway now hosts Laguna S 2.1 and Gemini 3.6 Flash with 1M context, letting you swap models in existing code without new SDKs. Evaluate coding harnesses as Lambda recommends before moving any model into production pipelines. Watch the Together AI YC cluster and OpenRouter pricing for cost signals on agent runs.

What to watch next

Watch for production numbers from the new Gemini Flash models on Vercel this week. Track early user reports on Buzz agent chat and the free Laguna tier on OpenRouter. Check Hugging Face for updates to the trending robotics and EEG models.

Harshs take

The day showed heavy concentration on model access through gateways rather than new capabilities. Most releases reduce friction for testing but still require builders to pick the right harness or context size to see gains. Hardware updates from NVIDIA target only the largest clusters, leaving smaller teams with the same token-cost tradeoffs.

The practical move is to pick one new model from the Gateway list, run a fixed agent task three times with different context settings, and record token use and output quality before adopting it further.

by Harsh Desai

Sources

Vendor launches

Hugging Face trending

Product Hunt picks

Other

Industry news

More AI news

Everything AI. One email.
Every Monday.

New tools. Model launches. Plugins. Repos. Tactics. The moves the sharpest builders are making right now, before everyone else.

No spam. Unsubscribe anytime.