OpenAI Agents on Vercel, Gemini Windows app, and new agent tools for builders
TL;DR
Vendors expanded agent hosting, model access, and inference options while new trending models and hardware integrations appeared across platforms.
What shipped
On 10 September 2026, releases focused on agent runtimes, model hosting, and inference improvements. Vercel led with multiple updates that tie OpenAI agents and GitHub Copilot into existing workflows. NVIDIA, Google, and smaller vendors added robotaxi tools, desktop apps, and specialized models.
Vendor launches
Vercel shipped the largest set of updates, integrating OpenAI Agents API and GitHub Copilot into its harness and sandbox layers. NVIDIA followed with robotaxi and physical AI announcements that target fleet-scale deployments. Google released a Windows Gemini app and expanded Dreambeans to more users.
- •OpenAI Agents on Vercel OpenAI now hosts long-running agent loops while Vercel supplies sandbox execution and queue handling for tool calls. Builders can connect sessions to isolated code environments without managing state themselves.
- •GitHub Copilot in AI SDK The new adapter lets applications swap GitHub Copilot into the same HarnessAgent interface used for other coding agents. Teams avoid rewriting code when testing different models.
- •Skild AI S1 model Skild AI released a foundation model that learns long-horizon factory tasks from one video. Manufacturers can update robots on changing lines without full reprogramming cycles.
- •Tako Search on AI Gateway The search tool runs free through September 30 and supplies grounded citations to any model on the gateway. Developers add web and graph data without new API keys.
- •Vercel Sandbox storage Every sandbox now ships with 64 GB of disk space instead of 32 GB. Teams working with large repos or data workloads gain headroom without resizing.
- •Gemini Windows app Google released a native Gemini desktop client that runs alongside other Windows tools. Users keep chat and generation open without browser tabs.
- •d-Matrix NVLink integration d-Matrix will connect its Raptor XPUs to NVIDIA NVLink Fusion for rack-scale inference. Chip users gain access to established scale-up networking without custom fabric work.
- •Vercel Sandbox regions Sandboxes now run in all 20 Vercel regions instead of four. Teams place execution closer to databases and meet data-residency rules with less latency.
Hugging Face trending
Six models climbed the trending list, covering text generation, image-text, and video tasks. Most run on the transformers or diffusers libraries and offer direct download or inference paths.
- •MiniCPM5-2B-GGUF openbmb released a 2B text-generation model in GGUF format that supports quick local inference through the Hub.
- •Spark-X2.5-4B-GGUF XHToken published a 4B text model in GGUF that runs on standard hardware with the gguf library.
- •Qwen-Drive-1.0-4B Qwen added a 4B image-text-to-text model aimed at driving scene understanding and captioning tasks.
- •Viggle-Animate Viggle released a video-to-video model that animates input clips while preserving subject motion.
- •Nex-N2.5-Pro nex-agi launched a text-generation model that supports fine-tuning and inference directly on the Hub.
- •DeepSeek-V4.1-Flash deepseek-ai released an image-text-to-text model optimized for fast multimodal responses.
Fal model gallery
GPT Image 2.5 Sunburst Edit: The model accepts targeted edit instructions and maintains subject and layout across multiple rounds. Users call it through Fal’s API for controlled image revision workflows.
Product Hunt picks
Four AI-related products launched, spanning UI inspection, agent workspaces, specialized small models, and hardware with AI features.
- •Modeinspect The tool offers 99 days of free credits for inspecting and designing product UI directly inside codebases.
- •hob The workspace manages an entire agent stack in one professional environment for teams running multiple agents.
- •Desert Ant Labs The company released small specialized models for speech, text, and vision tasks that fit on modest hardware.
Other
AWS, Databricks, OpenRouter, Together AI, and GitHub published workflow and inference updates. One historical note on Ethernet was excluded as unrelated to current AI tools.
- •Amazon Quick Automate RFI workflow AWS showed how to build an end-to-end RFI questionnaire using Quick Automate for repeatable document tasks.
- •Ling 3.0 Flash VL on OpenRouter InclusionAI made the free 262k-context vision-language model available at zero cost for testing multimodal routes.
- •Together AI open stack guide The post explains how to keep model, inference, gateway, and harness layers independent so new open models can be swapped in minutes.
- •Amazon Quick desktop app AWS released the desktop version of Amazon Quick for local access to automated ML workflows.
- •SageMaker prefix routing Amazon added prefix-aware routing to cut LLM latency on SageMaker Inference endpoints.
- •SageMaker HyperPod caching Amazon introduced model caching to reduce cold-start times on HyperPod inference clusters.
- •Bedrock Marengo 3.0 search Amazon enabled video and image search inside Bedrock Knowledge Bases using the Marengo 3.0 model.
- •GitHub Copilot app basics GitHub published a beginner guide for viewing diffs, running terminal commands, and previewing apps inside the Copilot desktop client.
Industry news
Anthropic, Meta, NVIDIA, and OpenAI made headlines on agent behavior, app rankings, growth forecasts, and subscription limits. Coverage also revisited AI risk discussions.
- •Anthropic CAPTCHA study Anthropic published findings that rogue agents struggle with CAPTCHAs in the same way humans do during web tasks.
- •NVIDIA 70 percent growth forecast Jensen Huang projected 70 percent revenue growth next year tied to broad AI infrastructure demand.
Replicate new models
p-video-2 on Replicate: prunaai released P-Video-2 at $0.025 per second as a quality-focused successor for video generation tasks.
What this means for you
For Vibe Builders: You can now connect OpenAI agent loops directly to Vercel sandboxes and run GitHub Copilot through the same harness interface without rewriting code. New small models on Hugging Face and Fal editing tools give you quick ways to test image and text workflows. Use the free Tako Search window and expanded sandbox regions to ship agent prototypes this week.
For Non-techies: Gemini now runs as a native Windows app so you can keep chats open while working in other programs. Dreambeans daily stories are open to every U.S. account and Meta’s Muse agent app ranks high for quick tasks. These releases move AI from browser tabs into tools you already use daily.
For Developers: Vercel’s 91 percent CDN metadata cut and 64 GB sandbox storage directly improve agent hosting performance at scale. NVIDIA’s NVLink Fusion and SageMaker prefix routing give concrete benchmarks for comparing inference stacks. Watch the independent layer approach from Together AI when swapping new open models into production pipelines.
What to watch next
Track OpenAI Pro subscription reopenings and any new agent harness adapters. Monitor Hugging Face for follow-up releases from the current trending models. Check region rollout notes from Vercel for latency-sensitive workloads.
Harsh’s take
The day’s releases cluster around integration layers rather than raw model capability jumps. Vercel’s volume of updates shows infrastructure vendors racing to lock in agent traffic, while NVIDIA continues to bundle physical AI into its existing platform story. The practical effect is more choice in hosting but added decision overhead for teams choosing between cloud sandboxes and local-first options.
A contrarian read is that many announcements repackage existing components with new marketing names. Builders should test one concrete integration, such as the OpenAI-Vercel agent path or the new Fal editor, against their current stack this week instead of collecting more vendor promises.
by Harsh Desai
Sources
Vendor launches
- •Build with OpenAI Agents API on Vercel
- •GitHub Copilot is now available in the AI SDK harness layer
- •Physical AI Takes the Wheel: How the World’s Robotaxi Leaders Are Building With NVIDIA Technologies
- •Skild AI Taps NVIDIA Physical AI to Teach Robots New Tasks From a Single Video
- •Dreambeans: Daily stories, brewed just for you, now available to all accounts in the U.S.
- •Tako Search is free on AI Gateway through September 30th
- •How we cut CDN metadata lookup latency by 91%
- •FastAPI frontends and static files served from the CDN
- •Vercel Sandbox now provides 64 GB of storage
- •The Gemini app is now available for Windows
- •Boots on the Ground: ‘WARDOGS’ Goes All Out on GeForce NOW at Early-Access Launch
- •d-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment
- •Exploring Creative Intelligence with London’s Southbank Centre
- •Vercel Sandbox is now available in all regions
Hugging Face trending
- •MiniCPM5-2B-GGUF by openbmb trends on HuggingFace
- •Spark-X2.5-4B-GGUF by XHToken trends on HuggingFace
- •Qwen-Drive-1.0-4B by Qwen trends on HuggingFace
- •Viggle-Animate by Viggle trends on HuggingFace
- •Nex-N2.5-Pro by nex-agi trends on HuggingFace
- •DeepSeek-V4.1-Flash by deepseek-ai trends on HuggingFace
Fal model gallery
Product Hunt picks
Other
- •Build an end-to-end RFI questionnaire workflow using Amazon Quick Automate
- •StoriesHow Kedde Used Lovable to Challenge Stockholm’s Zone-Based Logistics System
- •Improving Lakebase Postgres compute cache
- •inclusionAI: Ling 3.0 Flash VL (free) now available on OpenRouter (262k context, $0.00/M in, $0.00/M out)
- •InferenceThe Open Source AI StackA deep dive into the open model AI stack: model, inference, gateways and routers, harness, and tools, and how keeping each layer independent lets you swap in a new open model in minutes instead of rebuilding your workflow.
- •Amazon Quick is now generally available on desktop
- •Codeveloper of Ethernet Predecessor Dies at 91
- •Reduce LLM latency with prefix-aware routing on Amazon SageMaker Inference
- •Reduce inference cold starts on Amazon SageMaker HyperPod with model caching
- •Video and image search in Amazon Bedrock Knowledge Base using Marengo 3.0
- •GitHub Copilot app for Beginners: Using the diff, terminal, and browser
Industry news
- •Anthropic reveals rogue AI agents hate CAPTCHAs, just like you
- •Former Deepmind PR staffer says the lab once banned public discussion of AI extinction risk
- •Meta’s AI agent Muse is now the No. 2 app in the US
- •Jensen Huang explains why Nvidia will grow an astounding 70% next year
- •OpenAI puts Pro subscriptions on hold due to Astra demand
- •Is AI Actually Going to Kill Us All?
Replicate new models
More AI news
- Daily RoundupVercel CLI and v0 updates, Google AI tools, Apple iPhone AI features (agent memory and integrations)
Vercel expanded agent tooling and deployment security while Google and Apple rolled out practical AI features for everyday tasks and new hardware.
- Daily RoundupMercury 2.5 and Nex-N2.5 on OpenRouter, Meta Muse agent, and Vercel speed gains for agents
Model releases and agent tools expand options for builders while enterprise deals and routing improvements reduce friction for teams shipping AI products today.
- Daily RoundupHugging Face model wave, Fal H3 Turbo video, and Product Hunt AI agents
Hugging Face saw five models trend including text, image, video and speech tools while Fal released an upgraded text-to-video model and Product Hunt featured new agent-style apps for note-taking and local coding.