Gemini 3.7 Flash coding gains, Grok image model on Replicate, and agent marketplace moves
TL;DR
Google pushed Gemini 3.7 Flash for agents and coding while Replicate added image tools and Vercel expanded its marketplace and harness options.
What shipped
On 13 August 2026 several vendors released new models and integrations aimed at faster agent runs and easier image tasks. Google highlighted coding improvements in its latest Flash model while xAI and others added options on Replicate. The day also saw marketplace additions and Product Hunt tools focused on agent safety and video handling.
Replicate new models
Replicate added two image-focused models that let users generate and edit visuals without local hardware. xAI supplied the higher-resolution option while prunaai focused on virtual try-on for product photos. Both support direct API calls for quick integration into existing stacks.
- •Grok Imagine Image 2 xAI released Grok Imagine Image 2 on Replicate for text-to-image generation and editing up to 2K resolution. Teams can call it through the HTTP API to add image creation to marketing or design flows.
- •p-image-try-on prunaai released p-image-try-on on Replicate for virtual garment try-on on person photos. Vibe Builders can invoke it via API to test clothing visuals in product pages without new code.
Vendor launches
Vercel supplied the largest share of updates with GLM 5.2 access, Gemini 3.7 Flash pricing, Exa marketplace entry, and harness adapters. Google released Gemini 3.7 Flash for agent and coding tasks while also adding a Pixel notification feature. The changes center on lower-cost model runs and unified agent interfaces.
- •GLM 5.2 free tier Z.ai made GLM 5.2 free for eve agents through August 27 via Blackbox on AI Gateway. New agents default to the 1M-token model and can be started with a single CLI command.
- •Gemini 3.7 Flash release Google released Gemini 3.7 Flash as its strongest workhorse model for coding and agents. It handles design mocks to code and reduces failed tool-calling loops compared with prior Flash versions.
- •Gemini 3.7 Flash on AI Gateway Google listed Gemini 3.7 Flash on AI Gateway at 50 percent off until December 31. The discount applies to long agent runs where reliable multi-step execution matters most.
- •Exa on Vercel Marketplace Exa joined the Vercel Agent Marketplace as a native integration for search and research agents. Apps gain fresh context with one API key billed through the existing Vercel account.
- •Gemini Omni expert roundtable Google published an expert roundtable on Gemini Omni discussing model strengths. The session covers practical use cases for multimodal agent tasks.
- •ACP harness support Vercel added ACP-compatible harness support to the AI SDK via the new @ai-sdk/harness-acp package. Developers can now wrap any compliant harness instead of building one-off adapters.
- •Grok Build harness Vercel added Grok Build to the AI SDK harness layer through @ai-sdk/harness-grok-build. Existing code can switch to the new runtime without changes to the application layer.
Hugging Face trending
Three models trended on Hugging Face covering audio, text, and image-text tasks. MiniMaxAI supplied a text-to-audio option while deepseek-ai and meta-models added generation and multimodal models. All support download and fine-tuning through the Hub.
- •MiniMax-Music3 MiniMaxAI released MiniMax-Music3 as a trending text-to-audio model on Hugging Face. Builders can download it or run inference directly for music generation projects.
- •DeepSeek-V4-Pro-813 deepseek-ai released DeepSeek-V4-Pro-813 as a trending text-generation model on Hugging Face. It works with the transformers library for fine-tuning on coding or chat datasets.
- •Muse-Glimmer-30B-GGUF meta-models released Muse-Glimmer-30B-GGUF as a trending image-text-to-text model on Hugging Face. Users can fine-tune it for visual question answering or caption tasks.
Product Hunt picks
Six tools appeared on Product Hunt focused on agent control, video processing, chat portability, local browsing, AR creation, and AI search audits. The picks target both non-coders and developers who need quick safety or workflow fixes.
- •Phinq Phinq launched to stop AI agents before they cause damage. SMB owners can add it as a safety layer around automated tasks.
- •Qencode MCP Qencode released Qencode MCP so agents can transcode and process video. Teams gain direct video handling inside agent flows without separate services.
- •ThreadPort ThreadPort launched to move AI chats between ChatGPT, Claude, and Gemini in one click. Users keep conversation history when switching models mid-project.
- •Pickle Browser Pickle Browser released as a visible local browser for agents. Developers can watch and debug agent actions inside a contained window.
- •Kivicube Kivicube launched for no-code AR experiences built with AI. Vibe Builders can create product demos or training scenes without writing code.
- •AIO.GEO Protocol AIO.GEO Protocol released to audit AI search structure and test fixes. Site owners receive receipts after running dry-run optimizations.
Other
Posts covered software factory practices, Gemini 3.7 Flash pricing on OpenRouter, AI hallucinations, and smart routing for cost control. Augment Code and Databricks supplied the business-focused pieces while OpenRouter added batch pricing options.
- •Faster PR review loop Augment Code shared methods to shorten the path from PR to merge in software factories. The post focuses on review bottlenecks that slow agent-assisted coding.
- •Gemini 3.7 Flash on OpenRouter Google added Gemini 3.7 Flash to OpenRouter with 1,049k context at $0.38 per million input tokens. Builders can test the model in existing OpenRouter setups for agent runs.
- •AI hallucinations explained Databricks published a primer on AI hallucinations and why they occur. The piece gives practical checks for outputs in production chat systems.
- •Smart routing in Unity AI Gateway Databricks described smart routing that matches frontier quality at 30 percent lower cost per task. The method routes coding jobs across multiple models and harnesses.
- •Gemini 3.7 Flash batch on OpenRouter Google added Gemini 3.7 Flash batch mode to OpenRouter at half the standard price. Vibe Builders can run large offline jobs at $0.19 per million input tokens.
- •Gemini 3.7 Flash added on OpenRouter Google listed Gemini 3.7 Flash on OpenRouter for immediate testing by Vibe Builders and SMB owners. The model supports fast agentic workflows at listed rates.
Industry news
OpenAI hired a new CRO and partnered with IBM on enterprise training. Anthropic tested multi-agent behavior while Suno and Writer released product updates. Coverage also included pricing details on Gemini 3.7 Flash and a review of Zuckerberg's AI manifesto.
- •Anthropic agent turf war study Anthropic researchers ran AI agents on shared tasks and observed clashes and collusion. The findings question whether current safety tests cover multi-agent risks.
- •Suno Studio 2.0 chat feature Suno released Studio 2.0 with a chat feature that lets users create instruments via text. Premier subscribers gain unlimited MIDI import and 32-bit export.
- •Gemini 3.7 Flash pricing and benchmarks Google shipped Gemini 3.7 Flash three weeks after the prior version at half the price. The model beats Claude Sonnet 5 and GPT-5.6 Terra on coding benchmarks.
- •Writer new model and harness Writer released a post-trained GLM-5.2 variant with an upgraded harness to limit token spend. The system targets deployment-ready performance at lower cost.
What this means for you
For Vibe Builders: You can now call image generation and virtual try-on models directly on Replicate and test Gemini 3.7 Flash through OpenRouter or AI Gateway at discounted rates. Product Hunt tools like Phinq and ThreadPort give quick safety and chat portability without new code. These releases let you ship agent workflows and visual features faster by swapping in ready APIs instead of building from scratch.
For Non-techies: Gemini 3.7 Flash and GLM 5.2 free access lower the cost of running agents for daily tasks like research or content creation. New Replicate models handle product photos and try-on visuals while Suno Studio 2.0 adds chat-based music production. IBM and OpenAI enterprise moves signal more reliable options for small businesses that want supported AI tools.
For Developers: Vercel expanded the AI SDK harness layer with ACP support and Grok Build so you can swap runtimes without rewriting application code. Gemini 3.7 Flash shows gains on long tool-calling sequences and design-to-code tasks while OpenRouter batch pricing cuts offline job costs. Watch the multi-agent behavior findings from Anthropic before scaling coordinated agent systems in production.
What to watch next
Track Gemini 3.7 Flash adoption on OpenRouter and AI Gateway for price stability signals. Watch for follow-up posts on Anthropic multi-agent tests and any new harness adapters in the AI SDK. Check Replicate for additional image or audio models that reach trending status.
Harsh’s take
The day centered on incremental model drops and marketplace additions rather than fundamental shifts in reliability or cost structure. Many updates repackage existing capabilities under new pricing or harness layers without addressing core failure modes in long agent runs. The Anthropic multi-agent study stands out as a reminder that safety testing still lags behind deployment speed.
Builders should pick one concrete integration, such as the new ACP harness or a Replicate image endpoint, and run a timed task against their current stack this week. Measure actual loop failures and token spend before adding more tools.
by Harsh Desai
Sources
Replicate new models
Vendor launches
- •GLM 5.2 free for eve agents through August 27 via Blackbox on AI Gateway
- •Introducing Gemini 3.7 Flash
- •Gemini 3.7 Flash now available on AI Gateway for 50% off
- •Get updates while your Pixel 11 Pro is face down with HiLight.
- •Exa joins the Vercel Agent Marketplace
- •Omni experts share what excites them most about the model.
- •Use ACP-compatible harnesses with the AI SDK harness layer
- •Grok Build is now available in the AI SDK harness layer
- •One-click upgrade for deprecated Node.js versions
- •Inside the Vercel intern experience
Hugging Face trending
- •MiniMax-Music3 by MiniMaxAI trends on HuggingFace
- •DeepSeek-V4-Pro-813 by deepseek-ai trends on HuggingFace
- •Muse-Glimmer-30B-GGUF by meta-models trends on HuggingFace
Product Hunt picks
Other
- •Using the Cosmos software factory to better understand our business
- •The software factory needs a faster review loop: further optimizing the path from PR to merge
- •Google: Gemini 3.7 Flash now available on OpenRouter (1,049k context, $0.38/M in, $1.88/M out)
- •What are AI Hallucinations?
- •Smart Routing in Unity AI Gateway: Match frontier quality with 30%+ lower cost per task
- •Google: Gemini 3.7 Flash (batch) added on OpenRouter
Industry news
- •OpenAI hires new CRO as executive shake-up continues
- •IBM partners with OpenAI to bolster enterprise AI push
- •Anthropic set AI agents loose on the same task. They started a turf war.
- •Suno Studio 2.0's new chat feature lets you talk to your DAW like it's a bandmate
- •Gemini 3.7 Flash lands with coding gains and undercuts its three-week-old predecessor's price by 50%
- •Writer introduces new AI model and upgraded harness to contain token costs
- •Mark Zuckerberg’s AI Manifesto Is 6,500-Words, and Barely Says Anything
More AI news
- FeatureLovable adds option to turn off live preview for large projects
Project editors can now disable live preview to show the latest built version of an app instead of running a live dev server, improving performance.
- FeatureLovable generates Trust Centers for published apps
Lovable automatically generates a Trust Center security page at /.well-known/trust.html for every publicly published app along with a JSON twin listing observed security facts.
- Daily RoundupGrok 4.6 on AI Gateway, DeepSeek V4 Pro update, and new Replicate agent models
Vendors pushed model updates and agent infrastructure while device makers released new hardware with built-in AI features across the day.