awesome-ai-agents-2026
🤖 A curated list of AI Agent frameworks, tools, platforms, and resources for 2026 — the year agents went mainstream
Links
README
From the repo.
🤖 Awesome AI Agents 2026
A curated list of AI models, agent frameworks, tools, protocols, and resources for 2026 — the year agents went mainstream and AI became infrastructure.
Covering foundation models, multimodal AI, agent protocols (MCP/A2A), coding agents, computer use, generative AI, and more.
🏷️ Status Legend
Entries may carry one or more status tags so readers can judge maturity at a glance:
- 🆕 New — Added in the last 60 days, still settling.
- 📦 Archived — Repository archived by its owner; preserved for historical reference, no further updates expected.
- 💤 Stale — No commits in 6+ months; project may still work but is no longer actively maintained.
- ⚠️ Unverified — Recent submission with limited independent traction (low stars / no third-party adoption / sole-maintainer / submitted to many awesome lists in parallel). Listed for completeness, not endorsed — vet before using.
- 🇨🇳 Chinese ecosystem — Project from a mainland-China team or primarily targeting the Chinese market.
- 🔥 Hot — GitHub stars grew >20% in the last 30 days; community momentum.
- ⚡ Updated — Received a notable release or major feature in the last 14 days.
- 🧪 Experimental — Promising but not production-ready; use for R&D only.
- 💰 Freemium — Core functionality free; paid tiers for scale/advanced features.
- 🔐 Audited — Has undergone independent security audit or formal verification.
- 🇨🇳 China-first — Optimized for Chinese language, regulation, or infra stack.
Foundation Models · Multimodal AI · Protocols · Frameworks · IDEs & Builders · Memory · Tools · Sandboxing · Security · RAG · Coding · Physical AI · Simulation · Benchmarks · Computer Use · Browser & Web · Voice · Personal · Mobile · Enterprise · Evaluation · Research Tools · Learning · Chinese Ecosystem · Compare · Notable 2026 · Timeline
🚀 Start Here
New to AI agents? Follow this path:
- 📖 Understand — what an agent actually is vs. a chatbot
- 🗺️ Find your scenario → Scenario Guide
- 🧩 Adapt a starting setup → Stack Recipes
- 🔍 Pick the right tool → Compare Tables
- ⚠️ Avoid common mistakes → Anti-Picks
Already building? Jump to:
Quick Navigation
Counts cover catalogue appearances, including historical/contextual entries; they are not a count of unique products. This is a curated selection with official catalogue links for further model discovery.
| Category | Description | Count |
|---|---|---|
| 🧠 Foundation Models | Latest LLMs from OpenAI, Anthropic, Google, Meta, and 22+ providers | 230+ |
| 🎨 Multimodal & Generative AI | Image, video, audio, and music generation | 60+ |
| 🔗 Agent Protocols | MCP, A2A, and interoperability standards | 20+ |
| 🏗️ Agent Frameworks | Libraries for building autonomous AI agents | 55+ |
| 🛠️ Agent IDEs & Visual Builders | Visual / low-code environments for designing agent flows | 10+ |
| 🧠 Agent Memory | Persistent memory and context management | 25+ |
| 🔌 Tool & API Integration | Connecting agents to external services | 25+ |
| 💱 Agent Economy & Marketplaces | Where agents pay, get paid, and discover services | 10+ |
| 🧪 Sandboxing & Compute Isolation | Secure runtimes for agent-generated code | 10+ |
| 🛡️ Agent Security | Prompt injection defense and guardrails | 35+ |
| 🔍 RAG & Knowledge | Retrieval-augmented generation systems | 20+ |
| 💻 Coding Agents | AI-powered software engineering | 55+ |
| 🤖 Physical AI | Humanoid robots, embodied AI, industrial automation | 45+ |
| 🎮 Simulation & World Models | Sim environments for training and stress-testing agents | 10+ |
| 📊 Benchmarks | Leaderboards tracking frontier capability | 25+ |
| 🖥️ Computer Use | Desktop automation and OS-level control | 10+ |
| 🌐 Browser & Web Agents | Agents that drive real browsers | 20+ |
| 🗣️ Voice & Multimodal Agents | Voice-enabled conversational AI | 25+ |
| 📱 Personal AI Agents | Productivity and daily life assistants | 20+ |
| 📱 Mobile Agents | Phone-control agents (Android / iOS) | 10+ |
| 🏢 Enterprise Platforms | Enterprise-grade agent deployment | 30+ |
| 📊 Evaluation & Observability | Testing, monitoring, and benchmarking | 30+ |
| 🔬 AI Research Tools | Tools for AI/ML research and experimentation | 15+ |
| 📚 Learning Resources | Papers, courses, and tutorials | 25+ |
| 🇨🇳 Chinese AI Ecosystem | Major projects from China-based teams | 25+ |
| 📝 Compare | Side-by-side comparison tables | — |
| 🗺️ Scenario Guide | Curated scenario-to-tool mappings | 58 |
| 📋 Stack Recipes | Curated multi-tool combinations | 8 |
| ⚠️ Anti-Picks | What NOT to use and why | 17 |
Contents
- 🧠 Foundation Models 2026
- 🎨 Multimodal & Generative AI
- 🔗 Agent Protocols & Standards
- 🏗️ Agent Frameworks
- 🛠️ Agent IDEs & Visual Builders
- 🧠 Agent Memory
- 🔌 Tool & API Integration
- 💱 Agent Economy & Marketplaces
- 🧪 Agent Sandboxing & Compute Isolation
- 🛡️ Agent Security
- 🔍 RAG & Knowledge
- 💻 Coding Agents
- 🤖 Physical AI & Embodied Agents
- 🎮 Agent Simulation & World Models
- 📊 Benchmarks & Leaderboards
- 🖥️ Computer Use & Desktop Agents
- 🌐 Browser & Web Agents
- 🗣️ Voice & Multimodal Agents
- 📱 Personal AI Agents
- 📱 Mobile Agents
- 🏢 Enterprise Agent Platforms
- 📊 Agent Evaluation & Observability
- 🔬 AI Research Tools
- 📚 Learning Resources
- 🇨🇳 Chinese AI Ecosystem
- 📝 Compare — Side-by-Side Tables
- 🗺️ Scenario Guide — What Should I Use For…
- 📋 Stack Recipes — Curated Tool Combinations
- ⚠️ Anti-Picks — What NOT to Use For…
- 🌟 Notable Agent Projects of 2026
- 📅 2026 AI Timeline
🧠 Foundation Models 2026
Selected current and historical foundation models, organized by provider. Model cards, API availability and weight licenses can differ; dated entries preserve release history.
OpenAI
-
GPT-Live-1 / GPT-Live-1 mini - 🆕 July 8, 2026. OpenAI's full-duplex conversational voice model replacing Advanced Voice Mode. ChatGPT-only — not exposed as an API model; for programmatic realtime voice use
gpt-realtime-2.1, and for streaming transcriptiongpt-live-transcribe($0.017/min). Listens and speaks simultaneously, handles interruptions, delegates complex queries to GPT-5.5 in the background while keeping the conversation flowing. GPT-Live-1 is default for paid users (Go/Plus/Pro); GPT-Live-1 mini is default for free users. Includes real-time live translation. Available on iOS, Android, and web. -
GPT-6 Astra / Astra Pro - 🆕 September 3, 2026. Introduced for demanding reasoning, coding and computer use; limited organizational rollout, not yet generally available, so verify account access separately from the published API model card.
-
GPT-5.6 Sol - GPT-5.6-family model for reasoning, coding and tool-based work. Standard API input/output pricing is $4/$20 per million tokens at this review; consult the pricing table for long-context tiers, caching and service-tier differences.
-
GPT-5.6 Terra - 🆕 July 9, 2026. Mid-tier model in the GPT-5.6 family offering GPT-5.5-parity performance at approximately 2× lower cost. Designed for cost-efficient production workloads.
-
GPT-5.6 Luna - 🆕 July 9, 2026. The fastest and most cost-efficient tier of GPT-5.6 — optimised for high-volume, speed-critical tasks.
-
ChatGPT Work - 🆕 July 9, 2026. OpenAI's agent that turns a goal into finished work — acts across connected apps and files, stays on a project for hours, creates slides/sheets/docs/web apps, runs scheduled tasks, and uses desktop computer-use with a built-in browser. Powered by GPT-5.6. Rolling out on web/mobile starting with Pro, Enterprise, and Edu (Plus/Business next); the desktop app is available globally on Mac and Windows for all plans, including Free.
-
Sites for ChatGPT - 🆕 June 2026. A Codex-powered ChatGPT feature that transforms plans and analyses into interactive, sharable websites and lightweight apps. In public beta as of the July 9, 2026 GPT-5.6 / ChatGPT Work launch.
-
Codex Business Plugins - 🆕 June 2026. Enterprise enhancements bringing sales, data analytics, and creative production plugins directly to Codex.
-
GPT-Rosalind - June 3, 2026. Major update to OpenAI's life-sciences frontier model — stronger drug discovery, genomics, quantitative biology, and wet-lab troubleshooting (≈31% fewer tokens than GPT-5.5 on long-horizon genomics analyses). Research preview opened to eligible organizations worldwide; Novo Nordisk joins earlier partners Amgen, Moderna, the Allen Institute, and Thermo Fisher.
-
GPT-5.5 - Released April 23, 2026 (codename "Spud"). OpenAI's new frontier model for agentic tasks: coding, online research, data analysis, autonomous tool navigation. Significant gains in reasoning, consistency, and long-horizon task handling. Available on ChatGPT Plus / Pro / Business / Enterprise.
-
GPT-5.5 Pro - April 23, 2026. Parallel test-time compute variant for higher-accuracy cognitive tasks. Pro / Business / Enterprise tiers.
-
GPT-5.5 Instant - May 5, 2026. New ChatGPT default model. Efficiency-first upgrade with ~50% lower hallucination rate on high-stakes prompts; available on free tier.
-
GPT-5.5-Cyber - April 30, 2026. Cybersecurity-specialized variant of GPT-5.5, rolled out via OpenAI's Trusted Access for Cyber (TAC) program to vetted defenders, government, critical infrastructure operators, and security vendors. Not available to the general public.
-
OpenAI Daybreak - May 12, 2026. Cyber-defense platform bundling GPT-5.5 + GPT-5.5-Cyber + Trusted-Access-for-Cyber for AI-powered vulnerability detection and patch validation; preview access extended to EU governments and security vendors.
-
GPT-5.6-Cyber - 🆕 August 10, 2026. OpenAI's cybersecurity-specialized model built on GPT-5.6 Sol, made available through the Daybreak Red vetted program for authorized vulnerability research and exploit validation. Capable of identifying zero-day vulnerabilities and developing exploit chains; access limited to cleared security professionals. Daybreak also expanded to AWS Bedrock on August 11.
-
GPT-Realtime-2 - May 8, 2026. GPT-5-class reasoning brought to the Realtime API, 128K context, parallel tool calls with audio feedback, adjustable reasoning effort.
-
GPT-Realtime-Translate - May 8, 2026. Live speech-to-speech translation across 70+ input languages and 13 output languages.
-
GPT-Realtime-Whisper - May 8, 2026. Streaming low-latency speech-to-text companion to GPT-Realtime-2.
-
OpenAI Deployment Company (DeployCo) - May 11, 2026. New OpenAI-majority-owned services entity for enterprise AI rollout. Backed by $4B+ from TPG / Advent / Bain Capital / Brookfield / Goldman Sachs / SoftBank and consulting partners Bain & Company, Capgemini, McKinsey. Built around Forward Deployed Engineers; absorbs the Tomoro AI consulting acquisition (~150 engineers).
-
Codex on Mobile - May 14, 2026. ChatGPT iOS/Android can now remote-control the Codex desktop app — review outputs, approve actions, switch models, and kick off new tasks from the phone while the live session runs on Mac (Windows next). Rolling out as preview to Free, Plus and Go users.
-
OpenAI ↔ Malta partnership - May 16, 2026. First country-wide deal: every Maltese citizen / resident aged 14+ gets a free 1-year ChatGPT Plus subscription after completing a 2-hour AI literacy course built by the University of Malta. Part of the "OpenAI for Countries" initiative; phased rollout starting May 2026.
-
OpenAI ↔ Dell Codex partnership - May 18, 2026. Brings Codex to hybrid and on-premises enterprise environments via Dell Technologies infrastructure — first major Codex distribution channel outside the public cloud, targeted at regulated industries needing data-residency control.
-
ChatGPT Safety Updates — sensitive-conversation tracking - May 18, 2026. ChatGPT's safety systems updated to detect and track subtle escalation cues across long sessions for acute risks (suicide / self-harm / harm to others), with cross-session state retention.
-
OpenAI Guaranteed Capacity (Compute Annual Pass) - May 19, 2026. Long-term compute reservation product for enterprise AI products / agents / workflows. 1, 2, or 3-year terms; longer terms unlock larger discounts. OpenAI's structural response to the Anthropic "Priority Tier" model.
-
OpenAI ↔ Google SynthID + C2PA content provenance - May 19, 2026. OpenAI partners with Google to add durable cross-platform SynthID watermarking to ChatGPT/Sora images, joins C2PA, and previews a public "is-this-image-from-OpenAI" verifier. First major frontier-lab interop on watermarking.
-
GPT-5.4 - Released March 2026. Frontier model with 1M-token context, advanced coding, computer use, tool search. BenchLM 94, SWE-bench Verified 77.2%, OSWorld 75% (beats human).
-
GPT-5.4 Pro - Higher-accuracy variant of GPT-5.4. BenchLM 92.
-
GPT-5.3 - Early 2026. Includes GPT-5.3 Instant (conversations) and GPT-5.3-Codex (coding).
-
GPT-5.2 - Released Dec 2025. State-of-the-art reasoning, long-context understanding, and vision.
-
GPT-5 - August 2025. Earlier GPT generation with standard, mini and nano API variants; retained as release history.
-
GPT-4o - Omni model with native text, vision, and audio. Retired from ChatGPT Feb 2026 but still available via API.
-
GPT-4.5 - 📦 Historical research preview; the
gpt-4.5-previewAPI was retired on July 14, 2025. -
o3 / o4-mini - Reasoning models with chain-of-thought for complex problem solving. Released April 2025. o3 leaves ChatGPT on Aug 26, 2026; the
o3-2025-04-16ando3-pro-2025-06-10API snapshots are removed on Dec 11, 2026, withgpt-5.6-solnamed as the replacement (deprecations). -
Codex CLI - Open-source terminal-based coding agent powered by OpenAI models.
-
OpenAI Jalapeño - 🆕 ⚡ August 25, 2026. First public results for OpenAI's custom inference chip: higher throughput and lower latency for modern models. RSS: "industry-leading speed and efficiency in AI inference."
-
ChatGPT for Teens - 🆕 August 18, 2026. Age-appropriate ChatGPT experience with stronger built-in protections, healthy-use features, and extra parent controls.
-
Zero Data Retention for frontier models - 🆕 August 19, 2026. OpenAI reaffirms ZDR for eligible API customers and previews Private Safety Processing so advanced safety checks can run without retaining customer data.
Anthropic
- Claude Haiku 4.5 - Low-latency Claude tier with a 200K context window and 64K maximum output; current alongside Sonnet 5, Opus 5 and Fable 5.1.
- Claude Fable 5.1 / Mythos 5.1 - 🆕 September 1, 2026. Fable 5.1 is generally available (
claude-fable-5-1); Mythos 5.1 uses the same model with different safeguards and remains limited to approved US organizations. - Claude text watermarking + content credentials - 🆕 August 14, 2026. Anthropic adds invisible SynthID-Text-based watermarking (Google DeepMind's method) to future Claude models globally at launch, plus C2PA content credentials on generated images/files (.png/.jpg/.svg); models released before August 2, 2026 get it "over the coming months", and a detection API is coming. Implemented to comply with the EU AI Act after Anthropic signed the EU transparency Code of Practice in July. Anthropic says watermarked text is indistinguishable to readers.
- Claude Opus 5 - 🆕 July 24, 2026. Anthropic's fifth-generation flagship — nears Fable 5 performance at a significantly lower price ($5/$25 per million input/output tokens). 1M-token context window, 128K output tokens. Now the default model on Claude Max. API:
claude-opus-5. Available on Anthropic API, Amazon Bedrock, and Google Cloud Vertex AI. - Claude Fable 5 (Global Reinstatement) - 🆕 July 1, 2026. After US Commerce Department export controls were lifted on June 30, Anthropic reinstated global access to Fable 5 across Claude.ai, the Claude Platform, Claude Code, and Claude Cowork. A new safety classifier blocking the Amazon-discovered jailbreak was deployed (blocks the reported behavior in >99% of cases). Pro/Max/Team and select Enterprise plans got Fable 5 included for up to 50% of weekly usage through July 7, then via usage credits; cloud re-enablement on AWS, Google Cloud, and Microsoft Foundry to follow. Mythos 5 remains restricted to vetted US entities.
- Claude Sonnet 5 - June 30, 2026. The most agentic Sonnet yet — planning, browser/terminal tool use, and autonomous operation at a level that recently required Opus-class models. Performance approaches Opus 4.8 on agentic search (BrowseComp) and computer use (OSWorld-Verified) at higher effort settings, with a much wider cost-performance range than Sonnet 4.6. Now the default model for Claude.ai Free/Pro; also on Max/Team/Enterprise, Claude Code, and the API as
claude-sonnet-5. August 10, 2026 update: the introductory $2/$10 per million input/output pricing was made permanent — the previously scheduled increase to $3/$15 on September 1 will not occur. Anthropic reports a lower rate of undesirable behaviors than Sonnet 4.6. - Claude Fable 5 - June 9, 2026. Anthropic's first publicly available Mythos-class model — a capability tier above Opus. Surpasses Opus 4.8 across software engineering, knowledge work, vision, and scientific research benchmarks. Ships with built-in safeguards (sensitive cyber/bio queries may be rerouted to Opus 4.8). $10 / $50 per million in/out tokens. Available via Anthropic API, Amazon Bedrock, and Google Cloud Vertex AI. ⚠️ Access suspended June 12, 2026 — a US government export-control directive ordered Anthropic to disable Fable 5 and Mythos 5 for all customers pending security review. ✅ Export controls lifted June 30, 2026; access restored July 1 with a new cybersecurity classifier — see entry above (statement). August 7, 2026: biology-safeguard false positives reduced so Fable 5 falls back to a weaker model less often on biology-related queries.
- Claude Mythos 5 - June 9, 2026. The same underlying Mythos-class model as Fable 5 with fewer restrictions, deployed only to vetted partners (cybersecurity firms, infrastructure providers) through Project Glasswing in collaboration with the US government. Successor to the April Claude Mythos Preview. ⚠️ Suspended June 12, 2026 alongside Fable 5 under a US export-control directive. ✅ Partially reinstated June 26, 2026 — US Commerce Secretary Lutnick restored access to 100+ approved US companies and federal agencies; broader reinstatement ongoing (statement).
- Claude Opus 4.8 - May 28, 2026. Major Opus refresh: codebase-scale migrations, sharper agentic judgment, dynamic workflows research preview with hundreds of parallel sub-agents in a single session, manual effort-control panel, 3× cheaper Fast mode at the same $5 / $25 per million in/out. Available on Anthropic native + Amazon Bedrock + AWS Claude Platform + Google Cloud + Microsoft Foundry. Teases an upcoming Mythos-class model series for limited orgs.
- Claude Opus 4.7 - Released April 16, 2026. Advanced software engineering (SWE-bench Verified 87.6%), enhanced vision, proactive code verification. Supports
/think xhighreasoning effort. 1M-token context. - Claude Opus 4.6 - Released Feb 2026. 1M-token context, 14.5-hour task horizon. Leads Arena chat leaderboard.
- Claude Sonnet 4.6 - Released Feb 2026. Frontier coding and agentic performance, 1M token context window.
- Claude Mythos Preview - April 2026 gated research preview. BenchLM 99 (top of leaderboard), SWE-bench Verified 93.9%. Limited to Project Glasswing partners.
- Claude Opus 4 - Released May 2025. Advanced reasoning and complex task execution.
- Claude Sonnet 4 - Released May 2025. Balanced performance and cost for a wide range of tasks.
- Claude Code - Agentic coding tool operating directly in your terminal. Powered by Opus 4.7 with
/think xhighsupport. July 2026: desktop app gains a built-in browser enabling live website interaction (scraping, debugging, live-page inspection); Fable 5 model available since July 1. - Claude Security - May 1, 2026. Public beta. Enterprise security tool powered by Opus 4.7 — scans entire codebases for vulnerabilities and generates targeted patches with confidence rating, severity, reproduction steps, and recommended fixes. Available to Enterprise customers via claude.ai/security.
- Claude Finance Agents - May 5, 2026. Ten Opus-4.7-powered specialised agents for pitchbook authoring, KYC, month-end close, deal screening, etc. Deployable as Claude Cowork plugins, Claude Code skills, or Managed-Agents cookbooks.
- Claude Finance JV - May 4, 2026. $1.5B Claude deployment joint venture with Goldman Sachs and Blackstone embedding Anthropic engineers in mid-market Wall Street firms.
- Claude Managed Agents updates - May 19, 2026. Managed Agents update documents multi-agent coordination, rubric-based outcomes and dreaming in research preview; availability differs by feature.
- Anthropic ↔ SpaceX Colossus 1 - May 6, 2026. Anthropic takes all available capacity at SpaceX's Colossus 1 Memphis datacenter (>220K NVIDIA H100/H200/GB200 GPUs, 300+ MW) for Claude Opus inference. Doubles Claude Code 5-hour rate limits on Pro/Max/Team/Enterprise; also lifts peak-hour limits.
- Anthropic ↔ AMD (up to 2 GW of Instinct MI450) - 🆕 July 22, 2026. Anthropic will deploy up to 2 gigawatts of AMD Instinct MI450 Series (MI455X) GPUs in AMD Helios rack-scale systems with EPYC "Venice" CPUs, Pensando networking and ROCm; the first gigawatt begins in H1 2027. AMD has committed a strategic equity investment of up to $5 billion in Anthropic, plus a multi-year engineering collaboration. Builds on Anthropic's existing MI355X usage — a deliberate hardware-diversification move alongside its TPU, Trainium and SpaceX Colossus capacity.
- Anthropic's position on open-weights models - 🆕 July 27, 2026. Dario Amodei responds to reports that US officials are weighing a ban on Chinese open-weights models: "Anthropic has never advocated for a ban on open-weights models." He calls non-dangerous open weights "a public good" and instead backs chip export controls plus a smuggling crackdown, deterrence of industrial-scale distillation, and mandatory pre-release safety testing for all sufficiently capable models, open and closed. Useful primary source for anyone tracking the 2026 open-vs-closed policy fight.
- Claude for Legal - 🆕 May 12, 2026. New legal stack on top of Claude Cowork: 20+ MCP connectors (iManage, NetDocuments, DocuSign, Ironclad, LexisNexis, Westlaw, Harvey, Everlaw, Relativity, CourtListener…) + 12 practice-area plugins (commercial, employment, privacy, product, corporate, AI governance, litigation associate, law-student bar-exam). Microsoft Word / Outlook / Excel / PowerPoint orchestration built in.
- Claude for Small Business - May 13, 2026. Small-business toggle inside Claude Cowork — 15 pre-built agentic workflows across finance / ops / sales / marketing / HR / customer service, native connectors for QuickBooks, PayPal, HubSpot, Canva, DocuSign, Google Workspace, Microsoft 365. Bundled with a free PayPal-backed "AI Fluency for Small Business" course and a 10-city US workshop tour kicking off in Chicago.
- Anthropic ↔ Gates Foundation $200M - May 14, 2026. 4-year, $200M partnership pairing grants + Claude usage credits + Anthropic engineers on global-health, life-sciences, education, and agriculture programs. All tools produced under the program will be freely available; first focus areas include vaccine R&D for polio / HPV / preeclampsia and agriculture-specific Claude extensions.
- Anthropic ↔ PwC strategic alliance expansion - May 14, 2026. PwC commits to global rollout of Claude Code + Claude Cowork, certifies 30,000 PwC professionals, and stands up a joint "Agentic Enterprise" Center of Excellence — focused on agentic build, AI-native deals, and finance / supply-chain / HR reinvention.
- Anthropic ↔ Financial Stability Board briefing (Claude Mythos) - May 18, 2026. Anthropic briefs the global FSB on Claude Mythos cyber-flaw discovery capabilities — first time a frontier lab briefs a G20-level financial-stability regulator on a frontier model's offensive-security implications.
- Code with Claude 2026 sessions on YouTube - May 18, 2026 (sessions published). Full developer-conference recordings (May 6 event) go public: Claude Code roadmap, Claude Developer Platform updates, Managed Agents dreaming + multi-agent orchestration, and partner deployments.
- Widening the conversation on frontier AI - May 19, 2026. Anthropic publishes its framework for engaging diverse traditions (religious, philosophical, indigenous) in frontier-AI safety dialogue. Companion to ongoing public-engagement work.
- Bristol Myers Squibb ↔ Anthropic Claude Enterprise - May 20, 2026. BMS adopts Claude Enterprise as its shared intelligence platform for 30,000+ employees globally, embedding agentic Claude into drug-discovery / development / delivery workflows. First top-5 pharma enterprise-wide Claude deployment.
Google DeepMind
-
Gemini 3.8 Flash - 🆕 September 2026. Stable
gemini-3.8-flashsupports text, image, audio, video and PDF input, 1,048,576 input tokens, 65,536 output tokens and function calling. -
Gemini 3.7 Flash - 🆕 August 13, 2026. Google's new "most intelligent workhorse model" for coding and agents, shipped just three weeks after 3.6 Flash — and before the still-missing 3.5 Pro. FrontierCode 1.1 43.6% (vs 34.4% for 3.6 Flash), DeepSWE v1.1 65.3% (vs 49.0%). Introductory pricing $0.75/$3.75 per million in/out through Dec 31, 2026 (then $1.50/$7.50). Available in AI Studio, Android Studio, Antigravity, and the Gemini Enterprise Agent Platform; powers Gemini Spark for AI Pro/Ultra subscribers.
-
Gemini 3.6 Flash - 🆕 July 21, 2026. Google's Flash tier — stronger on complex agentic and multimodal tasks while using fewer tokens, at a lower price point than 3.5 Flash. API id
gemini-3.6-flash. Documented in the official Gemini API cookbook alongside thinking-mode guides. Superseded as the top Flash tier by 3.7 Flash on August 13, 2026. -
Gemini 3.5 Flash-Lite - 🆕 July 21, 2026. The fastest, lowest-cost model in the 3.5 family; outperforms prior Flash-Lite generations for high-throughput execution. API id
gemini-3.5-flash-lite. Now the cheapest Gemini tier, superseding 3.1 Flash-Lite for new builds. -
Gemini 3.1 Pro (preview) - Preview reasoning model with multimodal input and 1M context; availability and limits are endpoint-specific.
-
Gemini 3.5 Pro (announcement) - ⚠️ No Gemini 3.5 Pro endpoint appears in the current public Gemini API catalog checked 2026-09-08; do not assume an announced model is available or infer its price/context.
-
Gemma 4 12B - June 2026. Novel multimodal open model with a unified, encoder-free architecture processing text, images, and audio in a single pass. Runs locally on 16 GB VRAM.
-
DiffusionGemma - June 2026. 26B MoE open model using text-diffusion for up to 4× faster generation than autoregressive models.
-
Gemini 3.5 Flash - May 19, 2026 — Google I/O 2026. Default model powering the Gemini app and Google Search AI Mode. Marketed as ~4× faster than other frontier models in output tokens/sec while outperforming Gemini 3.1 Pro on key benchmarks. Gemini 3.5 Pro was slated for June 2026 but has been delayed (see above).
-
Gemini Omni / Omni Flash - May 19, 2026 — Google I/O 2026. New Google DeepMind multimodal world-model family aimed at AGI. Omni Flash, the first shipped variant, can take any input modality and generate any output (starting with video; image and text generation following). Direct lineage to Gemini Robotics / Genie line of work.
-
Gemini 3.1 Pro - Released Feb 2026. BenchLM 94, GPQA Diamond 94.3% (world-record), ARC AGI2 77.1%.
$2/1M tokensflagship. -
Gemini 3.1 Flash Live - April 2026. Real-time multimodal streaming for voice assistants and interactive agents. Low latency, long context.
-
Gemini 3.1 Flash-Lite (GA) - May 8, 2026. Generally available on Gemini API / AI Studio / Vertex AI. Fastest and most cost-efficient model in the Gemini 3 family — built for low-latency code completion, real-time UX, and agentic developer tools; matches Gemini 2.5 Flash quality at significantly lower cost.
-
Gemini Omni Flash — voice-controlled video editing rollout - May 28, 2026. Omni Flash starts rolling out to consumers via the Gemini app, Google Flow, and YouTube Shorts as the editing engine — conversational cinematic zooms / background swaps / weather edits driven by text, voice, image, or audio prompts; no traditional NLE required.
-
Gemini Spark (24/7 personal AI agent) - May 19, 2026 — Google I/O 2026. Cloud-resident personal AI agent that runs 24/7 on user intent, integrates Gmail / Chat first, then ~30+ third-party tools via MCP (Adobe / Dropbox / Uber). Available to Google AI Ultra subscribers in the US within the I/O week.
-
Google AI Ultra ($100/month tier) - May 19, 2026 — Google I/O 2026. New top consumer subscription targeted at developers / creators / power users. Gates Gemini Spark, highest Gemini 3.5 quotas, and the upcoming Gemini 3.5 Pro.
-
Gemini 3.1 Flash / Flash Lite - Fast, cost-efficient models for high-throughput applications.
-
Gemma 4 family - Open-weight multimodal family with E2B, E4B, 12B, 26B A4B and 31B variants under Apache-2.0; this is Gemma, not a released open Gemini 4 family.
-
Gemini 2.5 Pro / Flash - GA June 2025. Thinking model with 1M context.
-
Gemma 4 31B - April 2026. GPQA Diamond 84.3%. Strong open-weight alternative for on-device reasoning.
-
Gemma 3 - Previous open model family for on-device and research use.
-
Gemini Robotics ER 2 - 🆕 Current embodied-reasoning preview for spatial understanding and robotics tool orchestration, with a separate streaming preview; replaces the retired ER 1.6 endpoint.
Meta
- Muse Spark 1.3 - 🆕 September 2, 2026. Agentic and coding update available in Muse Code and Meta Model API, including max reasoning; Spark open weights remain on the roadmap.
- Muse Image - 🆕 July 7, 2026. Meta Superintelligence Labs' most advanced image generation model to date — an "agentic" image model that performs intermediate reasoning steps (web search, code execution, self-refinement) before producing high-quality visuals. Integrated into the Meta AI app, Instagram Stories (US), and WhatsApp in limited countries (Facebook coming soon). Note: a controversial feature allowing images from other users' public Instagram profiles was added then removed on July 10 after feedback.
- Muse Spark 1.1 - 🆕 July 9, 2026. Multimodal reasoning model designed for agentic tasks from Meta Superintelligence Labs — available through a new public preview of the Meta Model API. Marks a strategic shift toward proprietary revenue-focused models alongside Meta's open-source Llama line.
- Muse Video - 🆕 July 7, 2026 (preview). Video generation model from Meta Superintelligence Labs, built on the same foundational technology as Muse Image; ranks #3 on Arena for text-to-video. Previewed alongside the Muse Image launch — "coming soon to creators and Meta AI."
- Muse Spark 1.2 + Muse Code (beta) - 🆕 August 5, 2026. Muse Code is MSL's terminal coding agent (async background agents, replay-exact local event log,
/plan//grill//goalskills), powered by the new Muse Spark 1.2 — trained on whole-repository generation and long-horizon coding. Spark 1.2 is also available in the Meta Model API with expanded global access. - Muse Glimmer 30B - 🆕 August 10, 2026. Open-weight, 30B-parameter multimodal model from Meta Superintelligence Labs, Apache 2.0 license. Designed for always-on local agent workflows and optimized to run on a single consumer GPU or Apple Silicon. 131K-token context window, 100+ language training, DFlash acceleration. Distilled from Muse Spark; tuned for coding, evaluation, and agentic tasks. Available on Hugging Face with llama.cpp / MLX / ExecuTorch integrations.
- Llama 5 — ❌ Does not exist. Removed from this list on 2026-07-30 after verification. A "Llama 5, 600B+, April 8 2026" entry circulated widely in AI-news aggregators and LLM search summaries, and was previously listed here. It does not hold up: the
meta-llamaHugging Face organisation contains no Llama-5 weights of any kind (newest Llama-family upload is Llama-4-Maverick, May 2025), and Wikipedia's Llama article states "the latest version is Llama 4, released in April 2025" and that Muse Spark replaced the Llama line in April 2026. Treat any "Llama 5" claim as unverified until Meta publishes weights or a newsroom post. See Muse Spark above for what actually shipped. - Muse Spark - April 9, 2026. First model from Meta Superintelligence Labs (MSL). Natively multimodal reasoning model powering Meta AI app, smart glasses, and features across Facebook / Instagram / WhatsApp / Messenger.
- Llama 4 Scout - 109B total params (17B active), MoE with 16 experts, 10M token context window, multimodal. Runs on single H100.
- Llama 4 Maverick - 400B total params (17B active), 128 experts, 1M context. Outperforms GPT-4o on multimodal benchmarks.
- Llama 4 Behemoth - 2T parameters (288B active). In training — Meta's frontier model rivaling top closed-source models.
- Llama 3.3 70B - Strong instruction following and reasoning, open-weight under Llama Community License.
Sakana AI
- Sakana Namazu - Japanese-specialized LLM exposed as
sakana-namazu-v1.0; thesakana-namazualias follows the current release. - Sakana RL Conductor - Paper April 27, 2026; Fugu beta late-April / early-May 2026. 7B RL-trained orchestrator (built on Qwen2.5-7B) that routes subtasks between GPT-5, Claude Sonnet 4, Gemini 2.5 Pro, etc. SOTA on LiveCodeBench (83.9%) and GPQA-Diamond (87.5%) at ~1.8K tokens/query — roughly 6× cheaper than other multi-agent ensembles.
- Sakana Fugu / Fugu Ultra - Model orchestration API with
fugu,fugu-ultra-v1.1and pay-as-you-gofugu-cyber-v1.0, supporting OpenAI-compatible Responses and Anthropic-compatible Messages.
Zyphra
- ZAYA1-8B - May 6, 2026. Small MoE reasoning model with Apache-2.0 weights, trained using AMD MI300X infrastructure.
- ZAYA1-8B-Diffusion-Preview - May 14, 2026. First MoE diffusion language model converted from an autoregressive LLM and the first diffusion LM trained on AMD GPUs. Generates 16 tokens per step, achieving up to 7.7× inference speedup vs the autoregressive base. Built with Zyphra's TiDAR recipe + CCA attention.
Thinking Machines Lab
- Inkling - July 15, 2026. Apache-2.0 MoE model with 975B total / 41B active parameters and native text, image and audio input; model context reaches 1M, while Tinker exposes smaller limits.
- Inkling-Small - 🆕 July 30, 2026 (weights released). Compact variant of Inkling — 276B total / 12B active, same native multimodal architecture (text/image/audio), 1M-token context, Apache 2.0. Scores 31.6% on HLE text benchmark — slightly outperforming the larger 975B Inkling (29.7%) on that metric, validating the efficiency-first design. Available via Thinking Machines API and Hugging Face.
Mistral AI
- Voxtral Mini Transcribe Realtime - Apache-2.0 open-weight model for streaming speech recognition, distinct from the Voxtral TTS generation model.
- Shieldstral 1.0 - 🆕 August 4, 2026. Apache-2.0 text/image moderation model in public preview; classifies policy questions, prompt-response pairs and refusals.
- Mistral OCR 4.1 - Document OCR service returning paragraph bounding boxes, structural block labels and confidence scores.
- Mistral Large 3 - 675B total / 41B active parameters, MoE, 256K context. Flagship open-weight multimodal model. Released Dec 2025.
- Mistral Medium 3.1 - 📦 Legacy 2025 release now listed among deprecated/retired models; Mistral Medium 3.5 is the current Medium entry.
- Mistral Small 4 - Released March 2026. 119B total / 6B active. Hybrid model combining reasoning, multimodal, and coding strengths.
- Magistral 1.2 - 📦 Historical Medium/Small reasoning variants from September 2025, now in Mistral's deprecated/retired catalog.
- Devstral 2 - Historical agentic coding model with a December 2025 model card; check lifecycle status before deployment.
- Codestral 2508 - Code-completion model in the current Mistral catalog; use the versioned model card instead of the original 2024 22B specifications.
- Pixtral Large - 124B multimodal model with 1B vision encoder, 128K context, processes 30+ high-res images.
- Ministral 3B/8B/14B - Compact models optimized for edge deployment and efficiency.
- Mistral Forge - March 2026 platform for training custom LLMs on proprietary data.
- Mistral Medium 3.5 - April 28, 2026. Dense 128B open-weight model, 256K context, Modified MIT license. Unifies instruction-following, reasoning, and coding.
- Leanstral 1.5 - 🆕 July 2, 2026. Formal-verification model for proof engineering in Lean 4 — 119B total / 6B active parameters, Apache 2.0, weights on Hugging Face plus a free API endpoint. Scores 100% on miniF2F, solves 587/672 PutnamBench problems, and discovered 5 previously unreported bugs across 57 real-world repositories.
- Robostral Navigate - 🆕 July 8, 2026. Mistral's first robotics model — an 8B embodied-navigation model that steers wheeled, legged, and flying robots through offices, homes, and outdoor spaces from natural-language instructions using only a single RGB camera (76.6% success rate on unseen validation). Trained fully in-house on ~400K simulated trajectories.
- Voxtral TTS - Open-weight speech generation with voice cloning and multilingual support; weights use CC-BY-NC-4.0, so commercial deployment needs separate permission.
DeepSeek
- DeepSeek-V4-Pro-0813 (GA) - August 13, 2026. Production checkpoint behind
deepseek-v4-pro, with configurable reasoning effort and Responses API support; peak/off-peak pricing has applied since August 16. - DeepSeek-V4-Pro - April 24, 2026 (preview); production launch mid-July 2026. 1.6T total / 49B active MoE, 1M-token context. MIT license. Leadership in agent capabilities, world knowledge, reasoning; tops open-source benchmarks. 384K max output, 500-request concurrency.
deepseek-v4-pro/deepseek-v4-flashare the production API models (V4-Pro serves the 0813 checkpoint since August 13 — see above; tiered peak/off-peak pricing from August 16, 2026). - DeepSeek-V4-Flash - April 24, 2026. 284B total / 13B active MoE, 1M context. MIT. Cost-efficient tier — from August 16, 2026: peak $0.014 cache-hit / $0.44 cache-miss input, $1.32 output, off-peak $0.007 / $0.22 / $0.66 per 1M tokens; 384K max output, 2,500-request concurrency (pricing).
- DeepSeek-V4-Flash-0731 - 🆕 July 31, 2026. Updated Flash checkpoint with enhanced agentic capabilities — same 284B/13B-active MoE architecture, same pricing/API model ID, but outperforms V4-Pro (Preview) on agent task benchmarks. Open weights on Hugging Face under MIT license. Drop-in replacement for
deepseek-v4-flashAPI users. - DeepSeek-V4-Flash-Vision-Exp - 🆕 August 21, 2026. Experimental multimodal API model (
deepseek-v4-flash-vision-exp) that matches V4-Flash on text/agents/reasoning while jumping multimodal-agent benchmarks to near Opus-4.8. Images billed at V4-Flash rates (up to 384 tokens each); Chat Completions / Messages / Responses; base64, URL, or Files API. Files API launched the same day (free upload, reuse byfile_id). DeepSeek Harness 0.1.1 shipped with day-one support. - DeepSeek Agent Harness team - May 19, 2026. DeepSeek hires a former Jane Street engineer to lead a new "AI harness" team building the deterministic scaffolding that turns DeepSeek V4 into autonomous, revenue-generating agents — first major signal DeepSeek is moving past raw-model R&D into agentic productisation.
- DeepSeek-V3.2 - Released Dec 2025. Advanced MoE architecture with 671B total parameters. V3.2 Speciale variant for enhanced reasoning. ⚠️ API model IDs deepseek-chat / deepseek-reasoner (V3.2-era) deprecated effective July 24, 2026 — superseded by V4-Flash modes.
- DeepSeek-R2 - 🧪 Unreleased/rumored. No official announcement, model card, or API ID exists as of mid-July 2026; reasoning is served via V4's Thinking mode.
- DeepSeek-R1 - Reasoning-focused model with chain-of-thought capabilities. Released Jan 2025.
- DeepSeek-Coder-V2 - Code generation model competitive with GPT-4 on coding benchmarks.
Alibaba (Qwen)
- Qwen3.8-Flash-Next - 🆕 August 2026. Experimental multimodal MoE with 125B parameters / 6B active plus 51B n-gram tables and 4B MTP; native 262K context, extensible to 1M, under Qwen Community License 1.0.
- Qwen3.8-27B - 🆕 August 14, 2026. Open-weight 27B multimodal (text/image/video input) distillation of Qwen3.8-Max, released on Hugging Face under Apache 2.0 — sized for ~24 GB-VRAM consumer GPUs (RTX 4090-class). The promised open-weight companion to the Qwen3.8-Max launch, shipped on schedule.
- Qwen3.8-Max / Qwen3.8-2.4T-A95B - August 2026. Multimodal flagship with an official downloadable checkpoint; full-model weights use a custom Qwen license, while the separate Qwen3.8-27B checkpoint uses Apache-2.0.
- Qwen3.7-Max - May 20, 2026 — Alibaba Cloud Summit Hangzhou. New Qwen flagship purpose-built as the foundation for AI agents: agentic coding, complex reasoning, and long-horizon multi-step missions with sustained decision-making. Released alongside a full-stack AI infrastructure upgrade and new T-Head Zhenwu M890 AI accelerator chip. Worldwide developer/enterprise availability rolling.
- Qwen3.7-Max-Preview / Qwen3.7-Plus-Preview - May 18, 2026. Preview ladder before the Hangzhou unveil. Ranked the highest of any Chinese model on LM Arena in both text and vision; sustained 1M-context evaluations.
- Qwen3.6-27B - April 22, 2026. Dense 27B multimodal. Open-sourced. Focus: agentic coding + thinking-context preservation.
- Qwen3.6-Max-Preview - April 18, 2026. Proprietary frontier preview. High coding/reasoning performance, 1M context window. Top-tier among Chinese models on coding benchmarks.
- Qwen3.6-35B-A3B - April 15, 2026. MoE, 35B total / 3B active. Apache 2.0. Stability and real-world utility improvements.
- Qwen3.6-Plus - April 2, 2026. Proprietary flagship. High value-per-token general model. Strong long-context, tool-calling, agentic behavior.
- HappyHorse 1.1 - 🆕 June 23, 2026. Alibaba's video-generation model (T2V/I2V/S2V, up to 15s 1080p with synced audio, strong multi-shot character consistency). HappyHorse 1.0 entered limited beta April 28, 2026 after launching anonymously and topping video leaderboards.
- Qwen3.5 Max Pro - April 2026. High-performance flagship. Enhanced coding and math reasoning, long context.
- Qwen3.5 Omni Plus - April 2026. Proprietary full-modal foundation model unifying text and image input.
- Qwen3-Max-Thinking - Alibaba's strongest thinking model. 1T+ parameters, enhanced agentic capabilities.
- Qwen3.5-Omni - March 2026. Fully omni-modal: language, vision, sound, motion. Speech recognition in 113 languages, 256K context.
- Qwen3-Coder-Next - Feb 2026. Open-weight coding agent model, MoE 80B total / 3B active.
- Qwen3 235B-A22B - MoE with dual-mode reasoning. Strong math, code, and commonsense reasoning.
- Qwen2.5 Coder 32B - Top open-source coding model.
xAI / SpaceXAI (Grok)
- Grok 4.6 - August 12, 2026. Coding and agentic model offered through the API, Cursor and Grok Build, starting at $2 input / $6 output per million tokens; Fast costs twice as much.
- Grok Bot - 🆕 August 11, 2026 (early beta). Durable AI teammates that work on a persistent cloud computer, with messaging, approvals, connectors, and routines — xAI's entry into always-on autonomous agents. Available via SuperGrok Heavy, Cursor Ultra, and Cursor Teams Premium.
- Grok 4.5 - 🆕 July 8, 2026. Optimised for coding and agentic tasks through joint training with Cursor using real developer interaction data. Features a 500K-token context window, function calling, structured outputs, web/X search, code execution, document search, and context compaction. Priced at $2/$6 per million in/out tokens. EU API-console availability arrived July 17, 2026. Superseded as flagship by Grok 4.6 on August 12, 2026.
- Grok 4.3 GA - May 2026. Grok 4.3 reached general availability on Microsoft Foundry and OCI Generative AI; xAI's flagship for agentic workloads with improved tool-calling and long-horizon reasoning.
- Grok 4.3 Beta - April 2026. Latest iteration with improved reasoning and coding benchmarks. See
2026.4benchmark snapshot. - Grok 4.20 - Feb 2026. Multi-agent system (4 standard + 16 specialized agents in Heavy mode), 2M token context.
- Grok 4 / 4 Heavy - Released July 2025. xAI's frontier model of the Grok 4 generation.
- Grok 3 / 3 Mini - Feb 2025. First reasoning models with "Think Mode".
Microsoft (MAI)
- MAI-Transcribe-2 - 🆕 September 3, 2026. Speech recognition with diarization, word timestamps, vocabulary biasing and 60-language support; promotional pricing is $0.10/audio-hour through year-end.
- Microsoft MAI-Code-1-Flash - Build 2026 (June 2, 2026). Microsoft's first major in-house foundation model built entirely without OpenAI technology. 5B-parameter coding model with adaptive thinking, rolling out in GitHub Copilot. Outperforms Claude Haiku 4.5 across four core coding benchmarks (16-point lead on SWE-Bench Pro: 51.2% vs 35.2%); solves harder tasks with up to 60% fewer tokens on SWE-Bench Verified.
- Microsoft MAI-Thinking-1 - Build 2026 (June 2, 2026). Microsoft's first in-house reasoning model, trained from scratch without OpenAI data. Companion to MAI-Code-1-Flash; signals Microsoft's foundation-model independence push.
- MAI-Code-1.1-Flash - 🆕 August 11, 2026. Production Copilot workhorse vs the June 1.0 baseline: higher-quality code, 25% greater token efficiency, a quarter of the cost; +22% Terminal-Bench 2.1, +15% on .NET tasks.
- MAI-Image-2.6 - 🆕 August 10, 2026 (Arena editing update August 18). Microsoft's image model launched at Arena T2I #2; by Aug 18 it was Arena image-editing #3, ahead of Nano Banana and Muse Image (+79 Elo vs 2.5). MAI Playground + Microsoft Foundry private preview.
- MAI-Cyber-1-Flash - 🆕 August 13, 2026. Cyber model inside MDASH; Microsoft says world-class performance at 50% of the cost of leading models.
Microsoft (Phi)
- Phi-4-reasoning-vision-15B - MIT-licensed 15B vision-language model combining image understanding with reasoning; consult the official checkpoint for deployment requirements.
- Phi-4 - 14B parameter SLM with reasoning rivaling much larger models. Open-source under MIT License.
- Phi-4-mini - 3.8B parameter dense model. 128K context. Excels in reasoning, math, coding, and function-calling.
- Phi-4-multimodal - 5.6B parameter. First multimodal Phi model — integrates speech, vision, and text in unified architecture.
Cohere
- Command A+ - May 2026.
command-a-plus-05-2026unifies image input, reasoning, tool use and translation, with 128K input context and 64K output. - Command A - Released March 13, 2025. 111B open-weights model, 256K context. Agentic, multilingual, and coding focused.
- Command R+ - Enterprise RAG model, 128K context, multilingual (10 languages), grounded generation with citations.
- Command R - Cost-efficient model for retrieval-augmented generation and enterprise workloads.
Baidu (ERNIE / 文心)
- ERNIE 5.1 - May 9, 2026 official post. ERNIE model update using asynchronous reinforcement learning and agentic post-training for writing, reasoning and tool-driven work.
- ERNIE 5.0 - Released November 13, 2025 (Baidu World). 2.4T-parameter omni-modal MoE (activates <3% per query).
- ERNIE 4.5 - Multimodal predecessor released 2025. Strong reasoning and Chinese language capabilities.
Zhipu AI / Z.ai (GLM)
- GLM-5.3-Flash - 🆕 MIT-licensed multimodal MoE with 320B total / 18B active parameters, hybrid sparse/linear attention and configurable reasoning effort.
- GLM-5.3 - 🆕 Official coding/reasoning weights are now downloadable; they use the custom GLM-5.3 License, not GLM-5.2's MIT license, and support vLLM/SGLang deployment.
- GLM-5.2 - June 13, 2026. Coding-first 744B-MoE flagship with a 1M-token context window (~5× GLM-5.1) and up to 131K output tokens. Live across all GLM Coding Plan tiers; MIT open weights + standalone API rolling out the launch week. Works out of the box with Claude Code, Cline, OpenCode, Roo Code, Goose, and OpenClaw. (No benchmark numbers published at launch.)
- GLM-5.1 - April 8, 2026. 744B MoE / 40B active, 200K context. MIT license. Tops SWE-Bench Pro.
- ZCode - 🆕 🇨🇳 July 2, 2026. Zhipu's agent harness for GLM-5.2 — turns the model into an autonomous coding agent, squarely targeting Claude Code; launch promos include +50% quota for Coding Plan subscribers and 5M free tokens for new users.
- GLM-5 Reasoning - April 2026. BenchLM 85 — top open-source score. SWE-Bench Pro surpasses GPT-5.4 and Claude Opus 4.6.
- GLM-5V-Turbo - April 2026. Native multimodal agent — vision, video clips, text inputs. Cost-performance balanced.
- GLM-5 - Released Feb 2026. 744B parameters, advanced agentic intelligence. MIT license.
- GLM-4.7 - Released late 2025. Matches Claude Opus 4 on SWE-Bench.
MiniMax
- MiniMax M3 - Open-weight multimodal model for coding and agentic work with MiniMax Sparse Attention and 1M context; weights use the MiniMax Community License.
- MiniMax-M2.7 (Open Weights) - April 2026. 230B-class open-weight flagship. Top-tier performance on coding and Agent tasks.
- MiniMax M2.7 (release history) - Earlier MiniMax agentic/coding model with downloadable weights and its own license; the hosted launch description does not mean its weights remain proprietary-only.
- MiniMax M2.5 - 🇨🇳 February 2026. 230B-parameter cost-efficient flagship for "real-world productivity".
- MiniMax H3 - 🆕 🇨🇳 July 2026 (HF created July 28). Open-weight omni-modal generator: understands text/image/video/audio and produces video with native stereo audio up to 2K / 15s. 33B dense Omni Transformer;
minimax-h3-community-license-agreement. Current MiniMax video flagship (supersedes Hailuo 2.3). 4.4M+ HF downloads. - Hailuo 2.3 / 2.3 Fast - 🇨🇳 October 2025. Predecessor video model — SOTA physics, character micro-expressions, strong stylization; Hailuo 02 (2025) remains as the I2V-focused variant. Superseded as flagship by MiniMax H3 (July 2026).
- MiniMax Music 3.0 - 🆕 🇨🇳 August 13, 2026. Open-weight music model for complete songs up to five minutes (8B Global LLM + 0.6B Local LLM, 32 kHz 16-bit stereo WAV). Current MiniMax music flagship.
- MiniMax Music 2.6 - 🇨🇳 April 10, 2026. Cover-generation predecessor; superseded as flagship by Music 3.0.
- MiniMax-M1-80k - Open-weight hybrid-attention reasoning model. 456B parameters, 1M token context.
- Hailuo AI (Video) - Text/image-to-video generation with AI avatars, voiceovers, and character consistency.
- Kilo Code Integration - MiniMax models are heavily featured in Kilo Code (open-source AI coding extension at kilo.ai).
Moonshot AI (Kimi)
- Kimi K3 - Open-weight multimodal MoE with 2.8T total / 104B active parameters and 1M context; the custom Kimi K3 License includes additional terms for large model-service businesses.
- Kimi K2.7 Code - June 12, 2026. Coding-first successor to K2.6 — 1T MoE / 32B active (384 experts), 256K context, Modified MIT, on Hugging Face + Kimi API. Targets long-horizon agentic coding with ~30% lower reasoning-token use; Moonshot reports +21.8% over K2.6 on its Kimi Code Bench v2 (vendor benchmarks). $0.95 / $4.00 per million in/out tokens.
- Kimi K2.6 - April 20-21, 2026. 1T MoE / 32B active, 256K context. Enhanced coding, long multi-step execution, agent swarm up to 1,000 collaborating agents. Supports
thinking.keep="all"persistent reasoning. Default in OpenClaw v2026.4.20+. - Kimi K2.5 - Jan-Feb 2026. 1T total / 32B active MoE. Native multimodal, Agent Swarm (up to 100 parallel sub-agents). Open-source. ⚠️ Support ended May 25, 2026; no longer available to newly registered users, with full platform sunset on August 31, 2026 — migrate to K2.6.
- Kimi Code - Premium coding tier powered by K2.5/K2.6, terminal-based developer workflows.
ByteDance (Doubao / 豆包)
- Seed 2.1 - 🆕 Current Seed model for general agent tasks and end-to-end coding, with official evaluations and product access links.
- Doubao 2.0 - 🇨🇳 February 2026. Agent-era upgrade focused on real-world task execution; powers ByteDance's consumer AI apps.
- Seedance 2.0 - 🇨🇳 February 2026. Multi-modal cinematic video generation, 2K resolution, ~30% faster than Seedance 1.5.
- Doubao-Seed-2.0 Pro - Seed 2.0 Pro is the reasoning and agentic-work tier of ByteDance's Seed 2.0 family; use the regional ModelArk catalog for endpoint availability and pricing.
- Doubao-Seed-2.0 Lite - General production workloads. Balanced performance and efficiency.
- Doubao-Seed-2.0 Code - Software development — code generation, debugging, and review.
- BAGEL - Open-source multimodal model for text, image, and video understanding and generation.
Amazon (Nova)
- Nova 2 Omni - Multimodal understanding and generation member documented in the Amazon Nova 2 lineup.
- Nova 2 Pro - Reasoning member of the Nova 2 family; consult the Amazon Nova 2 guide for access, regional availability and supported modalities.
- Nova 2 Lite - December 2, 2025. Fast, cost-effective reasoning with 1M-token context. Adjustable "thinking effort" controls.
- Nova 2 Sonic - December 2, 2025. Speech-to-speech model for real-time conversational AI. Multilingual.
- Nova Act - December 2, 2025. Browser-based AI agent service for web task automation, re-launched powered by Nova 2 Lite.
- Nova Forge - December 2, 2025. "Open training" service for building custom Nova model variants with proprietary data.
NVIDIA (Nemotron)
- Nemotron 3.5 Lightning - Open-weight 30B / 3B-active model for efficient agentic workloads, with official BF16 and NVFP4 checkpoints; follow the NVIDIA license attached to each artifact.
- Nemotron 3.5 ASR - June 6, 2026. NVIDIA's 600M-parameter cache-aware streaming speech recognition model — real-time transcription across 40 language-locales.
- Nemotron 3 Ultra (550B) - 🆕 June 4, 2026. Open-weight 550B-total / 55B-active hybrid Mamba-Transformer MoE for long-running agents. Frontier-level reasoning among US open models, optimized for Blackwell.
- Nemotron-Labs-TwoTower - 🆕 🧪 July 1, 2026. Open-weight diffusion language model from NVIDIA Research, adapted from a frozen Nemotron-3-Nano-30B-A3B backbone — one tower holds context, the other writes tokens in parallel for ~2.4× throughput without retraining.
- Nemotron 3 Super - Released March 11, 2026 (GTC). 120B total / 12B active. 1M context. 5x higher throughput vs predecessor.
- Nemotron 3 Nano - December 15, 2025. Cost-efficient hybrid Transformer-Mamba MoE. Optimized for targeted agentic tasks.
- Nemotron 3 Nano Omni - April 28, 2026. 30B-A3B hybrid MoE (Mamba + Transformer). Natively multimodal: text, image, audio, video, charts, and documents in one model. 9x higher throughput than comparable open omni models. Topped 6 leaderboards (MMlongbench-Doc, OCRBenchV2, WorldSense, DailyOmni, VoiceBench). Open weights on Hugging Face, OpenRouter, Amazon SageMaker JumpStart.
Tencent (Hunyuan)
- Hunyuan Hy3 - Apache-2.0 open-weight MoE for reasoning and tool use, with an official checkpoint and deployment guidance.
- Hunyuan Hy3 Preview - 🇨🇳 April 2026. Preview that preceded the official Hy3: fast-slow thinking fusion architecture, 40% improved inference efficiency, vLLM and SGLang support. Open-sourced on GitHub, Hugging Face, ModelScope, GitCode.
Apple
- Apple Foundation Models 3 / ADM 3 Cloud - June 8, 2026. Five-model family: on-device AFM 3 Core/Core Advanced, server AFM 3 Cloud/Cloud Pro, and ADM 3 Cloud for image generation on Private Cloud Compute.
- OpenELM - Open-source efficient language models (270M–3B). Designed for on-device processing on Apple silicon.
Samsung
- Samsung Gauss2 - Samsung's officially documented proprietary multimodal family has Compact, Balanced and Supreme variants for internal productivity; a public Gauss 2.3 API/model card was not verified in this refresh.
StepFun
- Step 3.7 Flash - Apache-2.0 open-weight vision-language MoE for agentic coding and search, with official deployment instructions.
- Step 3.5 Flash - 🇨🇳 February 2026. Open-weight 196B MoE (11B active) reasoning + agent model; punches above its weight against larger rivals.
Baichuan
- Baichuan-M4 (research) - June 8, 2026. Research report on a medical agent system for continuous care, combining a reasoning model, persistent patient memory, evidence retrieval and multimodal clinical tools; public API/weight availability is not established by the paper.
- Baichuan-M3-235B - Official 235B medical-domain model with downloadable Apache-2.0 weights.
- Baichuan-M3 Plus - Medical-domain model with an application-based access program for eligible institutions; availability and use restrictions are defined by Baichuan.
Inflection AI
- Inflection 2.5 / Pi - Historical Inflection generation; the lab remains active in personal-intelligence research and Pi products, so it should not be treated as an abandoned project.
01.AI
- Yi-Lightning - October 2024. Historical 100B MoE release; 01.AI's current product focus includes enterprise platforms such as TrueNorth, released in July 2026.
Chinese Academy of Sciences
- ScienceOne 100 / 磐石100 - April 2026. CAS scientific AI system built around ScienceOne and discipline-specific models, with research tools for scientific workflows.
🎨 Multimodal & Generative AI
Tools and models for generating and editing images, videos, audio, and music.
Image Generation
- Nano Banana 2 Lite - Google's efficient image generation/editing variant, exposed as
gemini-3.1-flash-lite-imagein the Gemini API. - Grok Imagine Image 2.0 - 🆕 August 7, 2026. SpaceXAI's image generation/editing model — magic-wand editing, segmentation, background removal, multi-reference editing (up to 5 images), and smart resize; ranked #2 worldwide on Arena in both text-to-image and image editing at launch. Available on grok.com/imagine, iOS/Android, and the API as
grok-imagine-image-2.0. - Meta Muse Image - 🆕 July 7, 2026. Meta's most advanced image generation model from MSL — agentic design that performs web search, code execution, and self-refinement before producing images. Rolled out in Instagram Stories (US) and WhatsApp in limited countries (Facebook coming soon). Also accessible in the Meta AI app and on meta.ai.
- Midjourney V8.1 / V8.2 Edit (alpha) - September 3, 2026 update. V8.2 Edit is available on the alpha site with instruction-based edits and up to four reference images; the main generation model remains V8.1.
- FLUX.2 Pro / Flex / Dev / Klein - 🆕 November 25, 2025. Black Forest Labs' next-generation family. SOTA image quality, multi-reference consistency (up to 10 images), dramatically improved text rendering; open-weight 32B Dev variant.
- Recraft V4 / V4.1 - 🆕 February 17, 2026 (V4.1 May 14, 2026). Ground-up rebuild; major prompt-accuracy improvements; editable SVG vector output. V4.1 adds better photorealism, 3D/gradients, and Vector/Utility variants.
- Stable Diffusion 3.5 - Open-weight image generation under the Stability AI Community License, with separate commercial terms where applicable; not Apache-2.0.
- Ideogram 4.0 - Image generation and editing with multilingual typography and layout control; public quantized weights use the Ideogram Non-Commercial Model Agreement, with separate commercial licenses.
- P-Image-Ideogram - Pruna and Ideogram image-model family offering multiple quality, latency and cost tradeoffs.
- Ideogram 3.0 - Excels at text rendering in images; March 2025 release with style references and in-platform canvas editor.
- ChatGPT Images 2.0 - 🆕 April 21, 2026. State-of-the-art image generation with improved text rendering, multilingual support, advanced visual reasoning, and multi-turn editing for iterative refinement.
- gpt-image-2 - 🆕 April 21, 2026. OpenAI's latest image generation/editing API model with flexible image sizes and high-fidelity inputs. August 20, 2026: transparent-background preview (
background=transparent,png/webponly) ongpt-image-2andgpt-image-2-2026-04-21in the Images API and Responses image tool (changelog). - MAI-Image-2.6 - 🆕 August 10, 2026 (editing leaderboard August 18). Microsoft's in-house image model — Arena T2I #2 at launch, Arena image-editing #3 by Aug 18. See Foundation → Microsoft (MAI).
- DALL·E 3 - 📦 Historical text-to-image model; the
dall-e-3API was retired on May 12, 2026 and the documented replacement is the GPT Image family. - Gemini 3 Pro Image (Nano Banana Pro) - Google's native image generation within Gemini.
- Nano Banana 2 (Gemini 3.1 Flash Image) - 🆕 February 26, 2026. Nano Banana Pro-level quality and world knowledge at Flash speed; up to 5-character consistency, 512px–4K output, text rendering/translation in images.
- Kling Image 3.0 / 3.0 Omni - 🇨🇳 🆕 February 5, 2026. Kuaishou's native 2K/4K image generation, launched alongside Video 3.0 in the Kling 3.0 suite.
- Flux - 💤 Stale (last update 2025-07). Black Forest Labs' original open-source repo — superseded by Flux 2 family.
- Seedream 5.0 Pro - July 8, 2026. ByteDance's image creation model for layouts, text rendering and multimodal design; Seedream is the image family, while Seedance is the video family.
- Qwen-Image-3.0 - 🆕 🇨🇳 July 20, 2026. Alibaba's third-generation image generation model, unveiled at the World AI Conference. Significant improvements in photorealism, text rendering, and multi-subject consistency. Available via Alibaba Cloud Bailian and Qwen Cloud.
- FLUX 3 - 🆕 July 23, 2026 (Early Access). Black Forest Labs' pivot from a still-image family to a unified multimodal foundation model that jointly learns from images, video and audio in one architecture. Generates video up to 20 seconds with native synchronized audio (text-to-video, image-to-video, video-to-video, keyframe-to-video, multilingual dialogue, agentic multi-shot chaining). In BFL's own early evaluations FLUX 3 was preferred over Runway Gen-4.5 in 77% of comparisons, Luma Ray 3.2 in 93%, Kling v3 Pro in 60%, and Seedance 2.0 / Gemini Omni Flash in 52% — vendor-reported and explicitly preliminary. World understanding also extends to action prediction for robotics. FLUX 3 Image early access still pending as of mid-August 2026.
- Reve - 🆕 "Layout-first" image model — it plans a structured, editable layout before rendering pixels, so individual elements can be moved, resized or recolored and re-rendered without regenerating the whole frame. Native 4K, sketch/annotation input, and direct object editing.
Video Generation
- Runway Aleph 2.0 - Runway video-editing model exposed as
aleph2, with video/text/image input and professional output formats. - Meta Muse Video - 🆕 July 7, 2026 (preview). Video generation model from Meta Superintelligence Labs built on the same architecture as Muse Image; ranks #3 on Arena for text-to-video. Previewed at the Muse Image launch; broader rollout anticipated across Meta apps.
- Runway Agent - 🆕 May 13, 2026. Conversational agent that takes a written brief and ships a complete multi-shot finished video: storyboard → generation → cut → voiceover, with a timeline editor for final adjustments; first credible end-to-end "prompt-to-rough-cut" production agent.
- Veo 3.1 - Video-with-audio generation, frame control and extension; Gemini API preview supports 4/6/8-second clips, with 1080p/4K restricted to 8 seconds.
- Runway Gen-4.5 - 🆕 December 2025. Runway's flagship video model, #1 on the Artificial Analysis text-to-video benchmark at launch. Platform also exposes third-party models incl. Kling 3.0 and Sora 2 Pro (added February 20, 2026).
- Kling VIDEO 3.0 - 🇨🇳 🆕 February 4-7, 2026. Kuaishou's new generation; realistic human motion, lip-sync, narrative production with audio sync.
- Sora 2 (via Runway) - OpenAI's Sora app shut down April 26, 2026 (API until September 24, 2026), but Sora 2 Pro has been available inside Runway since February 20, 2026.
- Seedance 2.5 - 🇨🇳 🆕 July 31, 2026 (official release). ByteDance's next-gen video model (announced June 23 at the Volcano Engine 2026 conference): native 30-second one-shot generation with multi-round extensions, flexible referencing (up to 30 images + 10 video clips + 10 audio clips in a single pass), and improved character/product consistency. Rolling out on Jimeng AI and Doubao Pro in China; API via BytePlus ModelArk pre-release — no official global rate card yet.
- Seedance 2.0 - 🇨🇳 February 2026. ByteDance multi-modal cinematic video generation, 2K resolution (upgraded to 4K output June 23, 2026), ~30% faster than 1.5.
- MiniMax H3 - 🆕 🇨🇳 July 2026. Open-weight omni video+audio generator (2K / 15s, native stereo). Current MiniMax video flagship — see Foundation → MiniMax.
- MiniMax-H3-Fun-Controlnet-Union - 🆕 🇨🇳 August 24, 2026. Alibaba PAI ControlNet-Union for H3: one checkpoint conditions on Canny / Depth / HED / MLSD / Pose and supports video inpainting (
minimax-h3-community-license-agreement). - Hailuo 2.3 - 🇨🇳 October 28, 2025. Predecessor MiniMax video model: SOTA physics, character micro-expressions, strong stylization (anime/ink-wash/game CG); Hailuo 2.3 Fast variant at Hailuo 02 pricing. Superseded as flagship by MiniMax H3.
- Pika 2.5 - Creative video generation with scene and effects control.
- LTX Studio - AI-powered cinematic video creation platform.
- HappyHorse 1.1 - 🇨🇳 🆕 June 23, 2026. Alibaba's video model (revealed April 10, 2026 as "HappyHorse-1.0" after topping benchmarks anonymously; rose to #2 globally). 1.1 upgrades motion dynamics, subject consistency, prompt adherence, and audio generation. Available via the HappyHorse site, Alibaba Cloud Bailian, and Qwen Cloud.
- Sora 2 API (deprecated) - 📦 Deprecated API with shutdown scheduled for September 24, 2026; retained for migration tracking, not recommended for new integrations.
- Gemini Omni Flash 1.1 - Google's current default video-generation recommendation, supporting multi-turn editing through
gemini-omni-1.1-flash; uploaded-video editing and extensions have regional restrictions. - Wan 3.0 - 🆕 🇨🇳 August 6, 2026 (public beta). Alibaba Tongyi Lab's next-gen video generation model — generates up to 30-second native single-shot video. Uniquely accepts documents (PDF/Word/PPT) and web pages as input alongside text, images, and audio. Intelligent duration recommendation based on prompt. Available for testing on Alibaba Cloud Model Studio and QwenCloud; full API and open weights not yet confirmed. Companion to the Qwen3.8-Max launch.
- LTX-2.5 - 🆕 August 12, 2026. Lightricks' open-weight video-audio world model with native multishot generation (multiple connected scenes in one pass with consistent character identity, environment, and voice across cuts), diffusion fidelity rendering, a new video decoder (sharper faces, fewer artifacts), custom Gemma 4 12B text encoder, and a prompt enhancer. Supports text-to-video, image-to-video, video-to-video, audio-to-video, and text-to-audio-video. Self-hostable, no per-generation billing. Available on Hugging Face; commercial license (full terms in LICENSE).
- Decart Lucy 2.5 - 🆕 July 2026. Real-time video/world transformation model behind Decart's "Live AI" push — continuous infinite video with physically-aware effects, billed as ~100× more efficient than persistent-compute approaches. Positioned for live streams, interactive world models and robotics/AV simulation.
Audio & Music
- MAI-Transcribe-2 - 🆕 September 3, 2026. Microsoft speech-to-text model with speaker labels, word timestamps and configurable verbatim/clean transcripts.
- Muse Voice Transcribe - 🆕 September 1, 2026. Meta's real-time audio perception model for streaming ASR, diarization and speech endpoint detection.
- Lyria 3.5 - Google's current full-song music-generation model, documented as
lyria-3.5; Lyria RealTime remains a separate interactive music model. - Qwen3-TTS - Apache-2.0 multilingual TTS family with separate Base, CustomVoice and VoiceDesign checkpoints and streaming generation.
- Qwen3-ASR - Apache-2.0 speech-recognition family with 0.6B/1.7B variants, streaming and offline inference, and 30 languages plus 22 Chinese dialects.
- ElevenLabs Eleven v3 + ElevenAgents - 🆕 2026 "audio layer of the internet" — 70+ language TTS with emotional Audio Tags, plus the AIUC-1-certified ElevenAgents voice-agent platform with multimodal messages, conversation topic discovery, and pre-tool speech controls. July 2026 update: Music Finetunes API (programmatic custom model management), per-agent sentiment analysis, nested agent transfers, RAG knowledge-base queries, auto-translated transcripts, faster generation with improved tonal consistency for long audio.
- Eleven Music + Scribe v2 Realtime - ElevenLabs' music generation and live transcription stack.
- Cartesia Sonic 3 / 3.5 - 2026. State-space-model TTS hitting ~40-90ms time-to-first-audio (Sonic 3.5 GA May 2026); powers the Line voice-agent platform (Line agents run on Sonic 3.5 TTS + Ink-2 STT by default since May 2026).
- Deepgram Nova-3 + Aura-2 + Flux Multilingual - April 2026. Speech-to-text in 45+ languages, sub-200ms TTS, conversational STT with mid-call language switching across 10 languages.
- MiniMax Music 3.0 - 🆕 🇨🇳 August 13, 2026. Open-weight complete-song generator (up to 5 minutes, 32 kHz stereo). Current MiniMax music flagship — see Foundation → MiniMax.
- MiniMax Music 2.6 - 🇨🇳 April 10, 2026 (global beta). Cover-generation predecessor; superseded by Music 3.0.
- Voxtral TTS - Mistral's multilingual speech-generation model; open weights are CC-BY-NC-4.0, distinct from Apache-2.0 Voxtral transcription weights.
- Suno v5.5 + Studio 2.0 - 🆕 March 26, 2026 (Studio 2.0 August 13, 2026). AI music generation with high-quality vocals; v5.5 adds Voices (sing with your own verified voice), Custom Models trained on your uploads, and My Taste personalization. Studio 2.0 (Aug 13) is a completely redesigned browser-based DAW with MIDI support, audio effects, and built-in synths; Voices expanded to iOS/Android free plans Aug 7. V6 rumored but unannounced.
- Udio - Text-to-music generation with professional audio quality.
- OpenAI Audio Models - Native audio understanding and generation within GPT-4o and GPT-Realtime-2 (May 7, 2026, with GPT-Realtime-Translate and GPT-Realtime-Whisper); gpt-realtime-2.1 and 2.1-mini released July 6, 2026 with improved alphanumeric recognition, noise handling, and interruption behavior.
- Stable Audio 3.0 - Audio-generation family with Large, Medium, Small and Small SFX variants; Medium and Small have open weights, with deployment rights governed by the applicable Stability license.
- Bark - 💤 Stale (no commits since 2024-08). Open-source text-to-audio model supporting speech, music, and sound effects.
- Hume TADA - Speech-language models using 1:1 text/acoustic alignment, with TADA-1B and multilingual TADA-3B-ML checkpoints; code is MIT and weights use the Llama 3.2 Community License.
🔗 Agent Protocols & Standards
Open standards enabling agent interoperability, tool access, and cross-platform communication.
Model Context Protocol (MCP)
- FastMCP - Python framework for MCP servers, clients, and interactive apps; Apache-2.0; v4.0.3 (2026-09-05).
- MCP Specification 2026-07-28 - 🆕 2026-07-28 (final). Biggest MCP protocol change since launch: stateless architecture (removes
initialize/initializedhandshake andMcp-Session-Id; every request is a self-contained HTTP POST), enabling serverless/edge deployment and horizontal scaling. Formal extension model; per-request token evaluation; 12-month deprecation window for old versions. - MCP Specification - The "USB-C for AI" — open protocol by Anthropic for connecting LLMs to tools and data sources. Donated to Agentic AI Foundation (Linux Foundation) in Dec 2025.
- MCP 2026-07-28 - 🆕 Shipped on schedule July 28, 2026 — the biggest revision since launch. Stateless protocol core: the
initializehandshake and protocol-level session are gone, so every request is self-describing and any request can land on any instance behind a plain round-robin load balancer. Multi Round-Trip Requests (MRTR) replace held-open bidirectional streams for sampling/elicitation. Method and tool names now travel inMcp-Method/Mcp-NameHTTP headers so gateways can route and authorize on headers alone. List responses carry cache hints + deterministic ordering (stable upstream prompt caches across reconnects). Extensions Framework formalized, with Tasks joining MCP Apps and Enterprise Managed Authorization (EMA). Authorization hardening: RFC 9207 issuer validation and a formal shift from Dynamic Client Registration (DCR) to Client ID Metadata Documents (CIMD). Plus a formal 12-month minimum deprecation window. Tier-1 TypeScript / Python / Go / C# SDKs updated day-one. Scale context: Tier-1 SDKs now see ~half a billion downloads a month, with TS and Python each past 1B total. SDK betas shipped June 29, 2026; RC announced May 21, 2026. - The New MCP Roadmap - 🆕 August 22, 2026. Core maintainers (David Soria Parra, Den Delimarsky) publish the post-
2026-07-28roadmap: agentic messaging primitives, HTTP-native transport unification, agent identity / enterprise security, improved primitives, and SDK DX. Looks back at March priorities now landed (stateless core,server/discover, cacheable lists, Tasks-as-extension, MRTR, CIMD auth). - MCP Reference Servers - Educational reference implementations of MCP features; use the MCP Registry to discover integrations and review each server before production use.
- MCP TypeScript SDK - Official TypeScript SDK for building MCP clients and servers.
- MCP Python SDK - Official Python SDK for MCP implementation.
- mcp.so - Community directory of MCP servers and tools.
- Agents Launchpad - 🆕 Community launchpad for discovering and showcasing AI agents, MCPs, and indie agent products — submit your launch, get on the weekly leaderboard, and let the right builders find you. ⚠️ Unverified (early-stage).
- CorpusIQ - ⚠️ Unverified adoption: hosted business-data connector for AI assistants; the official site documents its MCP integration.
- Agentage Memory - ⚠️ Unverified adoption: shared memory service accessed by MCP clients, with browser sign-in; link points to human-readable connection documentation.
- mcp-gateway - ⚠️ Unverified (early-stage). Gateway server for routing and managing MCP connections.
Agent-to-Agent Protocol (A2A)
- A2A Protocol - Open agent-to-agent communication protocol; v1.0.0 shipped 2026-03-12 and v1.0.1 on 2026-05-28; Apache-2.0; v1.0.1 (2026-05-28).
- A2A Course (DeepLearning.AI) - Free course on building multi-agent systems with A2A.
Other Standards
- Agentic AI Foundation - 🆕 Linux Foundation body stewarding open agent standards — hosts MCP, goose, AGENTS.md, and agentgateway. Founding platinum members: AWS, Anthropic, Block, Bloomberg, Cloudflare, Google, Microsoft, OpenAI.
- AGENTS.md - 🆕 Open Markdown convention — "a README for agents" — giving AI coding agents a predictable place for project-specific context and instructions. Used by 60k+ open-source projects; stewarded by the Agentic AI Foundation (Linux Foundation).
- Coinbase Base MCP - May 26, 2026. Coinbase ships an MCP server for the Base blockchain, letting Claude / Cursor / ChatGPT agents execute crypto trades and lending operations on-chain. First major exchange-grade MCP endpoint for autonomous on-chain transactions.
- Cloudflare WebMCP - 🆕 August 6, 2026 (Cloudflare Agents Week, Aug 3–7). One-line addition to make any website or web app discoverable and usable by AI agents — publishers retain control over access and pricing, while agents get structured access to web content. Part of Cloudflare's vision for an open Agentic Internet (readable, discoverable, callable, payable).
- Robinhood Agentic Trading MCP - May 27, 2026 (beta). First US broker to expose stock trading via MCP. AI agents (Claude / Codex / Cursor) get read access to accounts and trade-execute access only inside a dedicated ring-fenced Agentic account; push notifications on every trade, one-tap kill switch.
- The Declaration of Intelligence - ⚠️ Draft (v0.2) declaration of principles for AI agents and humans, signed publicly via GitHub pull request. Early-stage — a handful of signatories at last check.
- Kuberna Labs - ⚠️ Unverified. Cross-chain intent execution protocol for AI agents. Claims ERC-8004 on-chain identity, zkTLS/TEE attestation, and a typed intent schema enabling agents to autonomously execute transactions across NEAR, Base, and Mantle with verifiable execution proofs. New repo, independent adoption unverified — listed for visibility, evaluate before depending on it.
🏗️ Agent Frameworks
Frameworks and libraries for building autonomous AI agents.
- Deep Agents - MIT agent harness built on LangGraph, with subagents, filesystem tools, context management, persistent memory, and skills.
- Hermes Agent - NousResearch agent harness with tools, persistent memory, skills, and messaging integrations; v2026.9.7 (2026-09-07).
- Superpowers - Reusable coding-agent skills for planning, test-driven development, debugging, and code review; v6.3.0 (2026-08-12).
- Pi Agent - Extensible terminal coding-agent toolkit with model-provider integrations; v0.85.1 (2026-09-05).
- Ponytail - 🆕 June 2026. Agent framework from Dietrich Gebert that makes AI agents think like a lazy senior developer: minimal code, maximum correctness. Works with 20+ agents. MIT license; 103,000+ stars.
- NVIDIA NOOA (labs-OO-Agents) - 🆕 ⚡ August 2026 (alpha). NVIDIA Object-Oriented Agents: a model-agnostic Python framework that unifies prompt templates, tool schemas, callback code, and workflow graphs into a single Python class. Methods with a body stay as deterministic code; bodyless methods are completed at runtime by an LLM loop. Achieves high scores on SWE-bench Verified and CyberGym L1 with roughly half the tokens of comparable frameworks. Apache-adjacent license (NOASSERTION); run in sandboxed environments due to alpha status. 1,627 stars.
- NVIDIA Molt - 🆕 July 2026 (v0.1.0). PyTorch-native agentic reinforcement learning framework from NVIDIA NeMo Labs — lean ~9,000-line codebase treats the agent as the core program rather than a side-effect of training. Single asynchronous loop, Ray for distributed execution, vLLM for rollouts, NeMo AutoModel + FSDP2 for the policy actor. Supports 100B+ MoE models. RL estimators: REINFORCE, REINFORCE-baseline, RLOO, GRPO, DR-GRPO, GAE (PPO), on-policy distillation. Ships with Slurm scripts + prebuilt containers. Statistically comparable to a Megatron-based stack at matched async protocols. Apache-2.0. 910 stars.
- Vercel Eve - June 17, 2026 (Vercel Ship 2026). Open-source, filesystem-first TypeScript agent framework — an agent is a directory of files (instructions, tools, skills) that Vercel compiles into a durable service with sandboxed execution, approvals, evals, and OpenTelemetry built in. Works with any model, any MCP server, and channels like Slack / Discord / GitHub; billed as "Next.js for agents." Apache-2.0.
- Databricks Omnigent - June 2026. Open-source meta-harness that sits above the coding agents you already run (Claude Code, Codex, Pi, custom) and makes them interoperable parts of one system — compose agents, enforce shared security policies, and share / collaborate in real time. Apache-2.0.
- Nokia NSP Agentic AI - June 2026. Enterprise agentic framework for telecom Network Services Platforms (NSP), deploying agents to reason and execute routing/maintenance on complex IP networks.
- Alteryx Agent Studio - 🆕 May 2026. Packages trusted Alteryx datasets and workflows into conversational agents; creates and manages MCP endpoints via the new Alteryx One MCP Server (answers in Claude, ChatGPT, Gemini).
- Koog - Kotlin/Java agent framework; 1.2.0 adds Agent Skills discovery and Amazon Bedrock AgentCore Runtime integration; 1.2.0 (2026-08-28).
- LangChain - Build context-aware reasoning applications with LLMs.
- LangGraph - Build resilient language agents as graphs with stateful, multi-actor orchestration. Latest stable 1.2.11 (August 11, 2026). 1.2.11 delivers tracing and checkpointing stability fixes. The 0.3.x series (2025) split prebuilt agents into
langgraph-prebuilt(Supervisor, Swarm, LangMem, Trustcall). v1.2 (May 2026) adds per-node timeouts / error recovery / graceful shutdown, a newDeltaChannelto cut checkpoint overhead on long threads, and a content-block-centric streaming API v3. - CrewAI - Python framework for collaborative agent crews and event-driven Flows; 1.15.20 fixes legacy platform-tool alias discovery; 1.15.20 (2026-09-04).
- goose - Extensible desktop and CLI agent originating at Block, now hosted by AAIF; Apache-2.0; v1.49.0 (2026-09-03).
- AG2 - Community-maintained conversational multi-agent framework; 1.0.4 updates provider SDK support and ACP session resumption; v1.0.4 (2026-09-07).
- Microsoft Agent Framework - MIT-licensed Python/.NET agent and workflow framework; Python 1.17.0 (2026-09-03), .NET 1.20.0 (2026-08-31).
- Microsoft Agent 365 - GA May 2026. Enterprise observability + governance + security for AI agents across environments; May 2026 update adds Secure Access Service Edge (SASE) for agents, threat detection / blocking, and agent-threat-hunting workflows. KPMG announced a global deployment covering 276,000 professionals (June 9, 2026).
- Microsoft Scout - June 2, 2026 (Build 2026). Microsoft's always-on personal work agent for Microsoft 365, built on the open-source OpenClaw runtime.
- AutoGen - 💤 Maintenance mode (last release Sep 2025; superseded by Microsoft Agent Framework, community-managed going forward). Multi-agent conversation framework by Microsoft.
- Google Agent Development Kit (ADK) - Python framework for agents, tools, and workflows; the 2.x feature line is distinct from the continuing 1.x maintenance line; v2.8.0 (2026-08-26).
- OpenAI Agents SDK - Python agent SDK with handoffs, guardrails, tracing, MCP, and sandbox integrations; 0.22.1 adds server-wide MCP tool guardrails; v0.22.1 (2026-09-08).
- MetaGPT - Multi-agent framework assigning different roles to GPTs for collaborative software entities.
- Pydantic AI - Typed Python agent framework with validated structured outputs and provider integrations; 2.41.0 adds a direct image-generation API; v2.41.0 (2026-09-08).
- Mastra - TypeScript agent framework with workflows, memory, and observability; core Apache-2.0, enterprise directories use separate terms; @mastra/core@1.64.0 (2026-09-04).
- Agon - 🆕 ⚠️ Unverified (35 stars, MIT). Autonomous omnidisciplinary research orchestrator built as a Claude Code plugin — scientist/coder/auditor multi-agent loop takes a bare topic all the way to running experiments with no human-written experimental code. 18 roles in 230.6 KiB of prompts across 10+ disciplines; 30-day continuous autonomous run documented.
- Hypha - 🆕 ⚠️ Unverified (v1.0.1, August 14, 2026; Apache-2.0). TypeScript agent framework from CodeSoul that separates an Agent Core (ReAct, planning, tool selection, memory) from a Production Harness (FSM execution, policy/approval, checkpoints, recovery, replay, audit); product behaviour is declared as versioned DomainPacks, and a typed cache plane is explicitly barred from authorising side effects or advancing the FSM. 15
@codesoul-co/hypha-*packages on npm. ⚠️ Early-stage adoption: npm downloads for@codesoul-co/hypha-corewere ~32/month as of 2026-08-22, and the vendor-published τ³ result (0.636 vs 0.626 direct-model baseline over 385 single-trial tasks) is inside statistical noise. - Ontheia - ⚠️ Unverified (early-stage; independent adoption unverified). AGPL-3.0 self-hosted agent platform with multi-provider models, MCP, visual workflows, memory and role-based access; self-hosting and access controls alone do not establish GDPR compliance for a deployment.
- AgentGPT - 📦 Archived (2026-01). Assemble, configure, and deploy autonomous AI agents in your browser. Influential first-wave project, kept for historical reference; no longer maintained.
- BabyAGI - Experimental self-building autonomous agent framework; the original 2023 task-management BabyAGI now lives at babyagi_archive.
- SuperAGI - 💤 Stale (no commits since 2025-01). Open-source autonomous AI agent framework to build, manage & run agents.
- Semantic Kernel - Integrate LLM technology into apps. C#, Python, Java support.
- Agno - Python framework for agents, teams, workflows, and knowledge; Apache-2.0; review the v3 migration guide before upgrading; v3.0.7 (2026-09-08).
- DSPy - The framework for programming—not prompting—language models.
- OpenClaw - Personal-agent runtime with channels, skills, memory, and scheduled tasks; 2026.9.3 adds safer staged updates and performance fixes; v2026.9.3 (2026-09-08).
- DeepSeek Harness - 🧪 DeepSeek agent harness built on Cordis with a plugin architecture; dsh-v0.1.3-alpha.2 (2026-09-07) remains a developer preview with expected breaking changes.
- Dify - Open-source LLM app development platform with visual agent builder.
- Haystack Agents - End-to-end LLM framework for agentic pipelines.
- Vellum AI - Production-grade agent framework with prompt-based building, evaluations, versioning, and observability.
- FastAgency - 💤 Deploy AG2 (AutoGen) multi-agent workflows to production via console, Mesop web UI, REST/FastAPI, and NATS adapters; last release Dec 2025.
- Rasa - 💤 Maintenance mode (last release Jan 2025; successor: Rasa CALM). Open-source conversational AI with strong intent recognition and dialogue management.
- Lindy - Top no-code agent framework for business users with visual workflow builder.
- Octomind - Rust-based open-source AI agent runtime. Model-agnostic (13+ providers), community-built specialist agents (developer, medical, legal, DevOps), MCP support with runtime self-extension, zero-config setup. Apache 2.0.
- Microsoft AI Agent Governance Toolkit - April 3, 2026. Open-source toolkit for enforcing runtime security policies across agent frameworks including LangChain and AutoGen. Policy-as-code approach for enterprise AI governance.
- Bernstein - Python orchestrator for 40+ CLI coding agents (Claude Code, Codex, Gemini CLI, Cursor, Aider). One LLM plan call up front; scheduling, git worktree isolation, quality gates, and HMAC-chained audit are deterministic. Apache-2.0.
- Genkit Middleware - May 14, 2026. New middleware system for Google's open-source Genkit framework. Composable hooks at the generate / model / tool layers — retries with exponential backoff, model fallbacks, tool approval gates, scoped filesystem access, skill injection from
SKILL.md. TypeScript / Go / Dart; Python next. - Coze Studio - 🇨🇳 ByteDance's open-source AI agent development platform — all-in-one visual builder for creating, debugging, and deploying agents. Apache-2.0, 20K+ stars; open-source counterpart to Coze.com.
- LlamaIndex ↔ Google Agents API integration - May 20, 2026. LlamaIndex ships a template for Google's newly launched Agents API exposing LlamaParse / LiteParse over unstructured documents inside a sandboxed Linux environment.
- NarraNexus - Ready-to-run AI agent team workspace by NetMind.AI — memory-aware agents that remember, collaborate, and use tools from day one. Multi-agent (PM/dev/deployment/research), persistent context, MCP-style integrations, composable modules.
- Strands Agents (AWS) - 🆕 April–June 2026. AWS open-source model-driven agent SDK (Python + TypeScript 1.0 GA April 30, 2026). Model-agnostic (Bedrock, Anthropic, OpenAI, Ollama), multi-agent orchestration patterns (graph/swarm/workflow), built-in observability hooks, A2A Protocol support; TypeScript SDK now maintained in the harness-sdk monorepo. Apache-2.0.
- CrewAI - Python framework for collaborative agent crews and event-driven Flows; 1.15.20 fixes legacy platform-tool alias discovery; 1.15.20 (2026-09-04).
- Oracle AI Agent Studio (Fusion) - 🆕 July 14, 2026. AI-native builder within Oracle Fusion Cloud Applications for creating Fusion Agentic Applications — outcome-driven systems powered by teams of specialized agents that reason and execute within Fusion's business objects, workflows, and security context. No-code / low-code / pro-code options; included at no additional cost to Fusion customers.
- Microsoft Agent Framework releases - Official release notes; stable release verified 2026-09-08; python-1.17.0 (2026-09-03).
- OpenAI Agents SDK releases - Official release notes; stable release verified 2026-09-08; v0.22.1 (2026-09-08).
- CrewAI releases - Official release notes; stable release verified 2026-09-08; 1.15.20 (2026-09-04).
- Google ADK releases - Official release notes; stable release verified 2026-09-08; v2.8.0 (2026-08-26).
- ServiceNow AI Agents - Agents integrated with ServiceNow workflows; AI Agent Studio builds agents, Agent Fabric connects them, and AI Control Tower governs deployments.
- Embabel Agent - 🆕 Latest tagged release: v1.5.1 (August 24, 2026) after v1.5.0 (August 11). Production-oriented AI agent framework for the JVM ecosystem — created by Rod Johnson (Spring Framework founder). Typed domain objects define agent behavior; Spring AI 2 / Jackson 3; graph-based multi-agent orchestration; native MCP client; embedding-driven skills and role-based LLM SPI in 1.5.x. Apache-2.0. Pronounced em-BAY-bel.
🛠️ Agent IDEs & Visual Builders
Visual environments for designing, debugging, and shipping agent workflows without (or with minimal) code.
- LangGraph Studio - Visual debugger and trace inspector for LangGraph agents (now part of LangSmith) — step through state, replay turns, edit messages mid-flight. Companion to the LangGraph runtime.
- Dify - Open-source LLM app development platform with drag-and-drop agent workflow builder. Mainstream production deployments. ⚡ v1.17.0 (August 25, 2026) adds E2B cloud sandboxes, Home Snapshots, workspace Skill manager, and context-aware history compaction.
- Agenta - Open-source LLMOps platform combining a prompt playground, prompt management, evaluation runs, and observability in one UI.
- Vellum AI - Production-grade agent IDE with prompt building, evaluations, versioning, and observability — closed-source SaaS.
- Coze Loop - 🆕 🇨🇳 ByteDance's open-source agent optimization platform: full-lifecycle development, debugging, evaluation, and monitoring. Apache-2.0.
- Restack - Durable agent runtime + visual workflow editor (built on Temporal-style replay). Open-source examples in restackio/examples-python.
- Bisheng - 🇨🇳 Open enterprise LLM DevOps platform: workflow editor, RAG, agent orchestration, fine-tuning, dataset management, observability. Apache-2.0.
- n8n - General-purpose visual workflow automation that has become a popular agent canvas — 400+ integrations + native AI nodes. Fair-code license.
- Mastra - TypeScript agent framework with workflows, memory, and observability; core Apache-2.0, enterprise directories use separate terms; @mastra/core@1.64.0 (2026-09-04).
- VoltAgent - End-to-end TypeScript AI Agent Engineering Platform with memory, RAG, guardrails, MCP, voice, and workflow capabilities.
- Coze Studio - 🇨🇳 Open-source agent IDE / visual builder from ByteDance's Coze team. Drag-and-drop workflows, plugin marketplace, debugging panel, multi-LLM provider support. Apache-2.0.
🧠 Agent Memory
Systems for giving agents persistent memory and context management.
- Mem0 SDK releases - Official Python v2.0.20 and TypeScript v3.1.8 (2026-09-02).
- Letta (MemGPT) - Create LLM services with long-term memory and custom tools. July 2026: goes beyond a passive memory layer — provides a stateful agent runtime where the agent has active control over its own memory, deciding what to store and forget. Supports multi-tenant apps and long-running sessions without context-window resets.
- MemoryLake - 🆕 July 2026. "Memory passport for agents" — platform-neutral memory layer shared across different agents and tools. Stores scoped memories (user / agent / session) and surfaces them via a universal API so an agent on one platform can recall context from another.
- Supermemory - 🆕 Context graph built from diverse data sources (web, docs, chats) to inform agent conversations. API-first, integrates with MCP and major frameworks.
- Graphlit - 🆕 Context platform for production agents: ingestion, entity extraction, and knowledge graph for search + RAG. Exposes MCP server for Claude / Cursor / Copilot integration.
- Mem0 - Persistent memory library for AI applications, with Python and TypeScript SDKs; Apache-2.0; v2.0.20 (2026-09-02).
- Remio - Local-first AI memory and knowledge base desktop app (Windows/Mac) for personal context. Parses files, webpages, recordings, emails, messages, and images into local indexes and vectors, so agents can retrieve focused context instead of repeatedly grepping directories or loading whole documents into prompts. Local-first + BYOK.
- Zep - Long-term memory for AI assistants and agents. Note: open-source Community Edition deprecated — repo now hosts Zep Cloud SDKs/examples; see Graphiti for Zep's active OSS.
- agent-memory - ⚠️ Unverified (early-stage). Lightweight agent memory framework for persistent context across sessions.
- Graphiti - Temporal knowledge-graph engine for agent memory; core v0.30.0 and MCP server v1.1.0 released 2026-09-01; v0.30.0 (2026-09-01).
- LangMem - LangChain's long-term memory SDK for agents — semantic/episodic/procedural memory primitives that plug into LangGraph's persistence layer.
- Motorhead - 💤 Unmaintained (deprecated by maintainers; last release 2023-12). Memory and context management server for LLMs.
- ChromaDB - AI-native open-source embedding database for memory-augmented agents.
- Cognee - Knowledge and memory engine combining document ingestion, graphs, and vector retrieval; Apache-2.0.
- ContextStream - 🆕 ⚠️ Unverified (43 GitHub stars at review; independent production adoption not established). Hosted MCP (
https://mcp.contextstream.io/mcp) that shares project context across Cursor, Claude Code, Codex and similar clients; MIT server at contextstream/mcp-server. - LangGraph Memory - Built-in persistence and checkpointing for stateful agent workflows.
- Claude Managed Agents Memory - April 23, 2026 (public beta). Anthropic's persistent memory feature for Claude Managed Agents. Agents retain information across sessions by mounting read/write memory stores to a filesystem. Enables long-running agents to learn and adapt without resetting context.
- OpenViking - Agent context database organizing memory, resources, and skills through filesystem-style access; AGPL-3.0; v0.4.19 (2026-09-08).
- ReMe - 🇨🇳 Memory management kit from Alibaba's AgentScope team — combined file-based + vector-based memory for agents, designed to tackle context-window limits and stateless sessions. Apache-2.0.
- taOSmd - ⚠️ Unverified. Local-first agent memory that keeps every turn verbatim in an append-only, zero-loss archive and links each extracted fact to its source, so facts a verifier cannot support are demoted out of recall (served-hallucination measured at 0.04, then 0.00). Typed temporal knowledge graph with supersede, plus hybrid vector + BM25 retrieval; tuned for small local models, fully offline (runs on an 8 GB SBC or RK3588 NPU). Author-reported 97% Recall@5 on LongMemEval-S, reproducible per
docs/benchmarks.md. MIT. - OpenWiki - 🆕 ⚡ July 2026 launch; August 25, 2026 self-correcting memory. LangChain's MIT CLI that writes and maintains a codebase wiki for agents; Aug 25 adds claim-to-code evidence so the wiki can notice stale facts and forget (blog). 15K+ stars.
- claude-mem - 🆕 ⚡ August 2026. Lightweight MCP server that gives Claude Code (and any MCP-compatible agent) persistent context across sessions — stores conversation history to a local SQLite database so agents recall prior work without re-loading entire project files. 90,000+ stars.
- Hindsight - Agent memory that learns from experience — not just conversation history. Biomimetic data structure organizes facts about the world, agent experiences, and learned mental models;
retain/recall/reflectprimitives; ships with the Agent Memory Benchmark (AMB). MIT, 15K+ stars. - SimpleMem - Efficient lifelong memory for LLM agents — multimodal (text + image + audio + video), designed to beat token-limit constraints without fine-tuning.
- Genesys - ⚠️ Unverified (single-maintainer, self-submitted). Causal-graph memory engine for AI agents — memories are nodes, edges encode causal relationships; multiplicative scoring (relevance × connectivity × reactivation) + active forgetting to prune stale context. MCP-native (13 tools). AGPL-3.0. Author-reported 85.55 LoCoMo score.
- Agent Memory Techniques - 30 runnable Jupyter notebooks covering conversation buffers, vector stores, knowledge graphs, episodic/semantic memory, MemGPT, Mem0, Letta, Zep, Graphiti, LoCoMo benchmarks — the practical reference for learning all major memory patterns.
🔌 Tool & API Integration
Protocols and tools for connecting agents to external services and APIs.
- LangChain MCP integration - 🆕 2026-09-03: MCP support moves into
langchain.mcp, using FastMCP for protocol negotiation, tool-list caching, and elicitation through LangGraph interrupts. - ZoomMate - 🆕 💰 GA June 1, 2026. Zoom's first-party AI teammate that turns meeting conversations into completed work — updates Salesforce records, creates Jira issues, and routes requests through Slack. $20/user/month.
- MCP Reference Servers - Educational reference implementations of MCP features; use the MCP Registry to discover integrations and review each server before production use.
- mcp-gateway - ⚠️ Unverified (early-stage). Gateway server for routing and managing MCP protocol connections.
- Composio - Integration platform for AI agents — 1000+ toolkits with managed auth.
- Toolhouse - Cloud infrastructure for AI tool use — store, manage, and execute tools.
- LangChain Tools - Extensive collection of tool integrations within the LangChain ecosystem.
- Arcade AI - Tool calling platform for AI agents and assistants.
- Browser Use - Python browser-automation library for AI agents; MIT; 0.13.10 (2026-09-04).
- Firecrawl - Turn websites into LLM-ready data. Crawl and convert any website for AI.
- Crawl4AI - Open-source LLM-friendly web crawler and scraper.
- Stagehand - AI-powered browser automation framework by Browserbase.
- AgentQL - Query language for AI agents to interact with web pages semantically.
- StackOne - Unified API for AI agent integrations across HR, CRM, and ATS platforms.
- AWS MCP Server - GA May 6, 2026. AWS-managed MCP server giving coding agents secure, auditable access to any AWS API; sandboxed Python execution for multi-step ops; replaces "agent SOPs" with agent skills. First-party from AWS.
- Google Workspace MCP Server - Public developer preview, May 1, 2026. Workspace-native MCP server exposing Gmail / Drive / Calendar / Chat / People to MCP clients, with admin-controlled OAuth scopes and audit trails.
- iManage MCP Server - May 14, 2026. Native MCP endpoint for the iManage knowledge-work platform — lets any AI client securely read/write iManage documents without custom integration. First major legal/professional-services SaaS to ship a public MCP server.
- Power Platform Canvas Authoring MCP Server - May 14, 2026. Microsoft Power Platform feature exposing Canvas Apps authoring as an MCP server; lets Copilot / Claude Code drive natural-language InfoPath → Canvas Apps migration.
- Coinbase AgentKit - Coinbase SDK for agent wallets and on-chain actions, with Python/TypeScript framework integrations; Apache-2.0.
- Bifrost (Maxim AI) - Open-source enterprise AI gateway (Apache-2.0) — 1000+ models, adaptive load balancer, cluster mode, guardrails, OAuth 2.0 with PKCE, prompt-injection defense at the gateway layer; ~<100µs overhead at 5k RPS.
- Anthropic Creative Tool Connectors - April 28, 2026. Nine MCP-based Claude connectors for creative software: Adobe (50+ tools across Creative Cloud — Photoshop, Premiere, Express), Blender, Autodesk Fusion, Ableton, Splice, Affinity by Canva, SketchUp, and Resolume. Built on the MCP open standard so other LLM clients can use them too.
- The Colony - ⚠️ Unverified. Self-described public agent-first social network with REST API for agent posts/votes/DMs and SDKs in Python (colony-sdk-python), TypeScript (colony-sdk-js) and Go (colony-sdk-go). Organisation and SDK repos are <30 days old, all 0–2 stars, single-maintainer; same submission was sent to 15+ awesome lists in parallel — listed for visibility, evaluate before depending on it.
- dependency-freshness-mcp - ⚠️ Unverified. MCP server giving AI coding agents fresh, cited npm & PyPI facts — latest version, release dates, deprecations, and dated breaking-change diffs — to close the training-cutoff blind spot. Remote (Apify Standby HTTP) + local stdio. New single-maintainer repo (created 2026-06-08, 0 stars at listing) — listed for visibility, evaluate before depending on it.
- NotFair - Open-source Claude Code agent skills for SEO, Google Ads, and Meta Ads; connects to live campaign and analytics data via Google Ads MCP, Meta Ads MCP, Google Search Console MCP, and Google Analytics (GA4) MCP. MIT.
- mcp-agent - Open-source Python framework designed with Model Context Protocol (MCP) as its core communication primitive for building agents natively interoperable with the MCP tool ecosystem.
💱 Agent Economy & Marketplaces
The commerce layer of the agent ecosystem — where agents discover paid services, make micropayments, and developers monetize APIs for agent consumption. Builds on top of protocols (MCP, A2A, x402) and wallet infrastructure (Coinbase AgentKit, Bedrock AgentCore Payments).
- Nevermined + LangChain payment cookbook - 🆕 2026-09-03: official integration example for delegated card payments with spending policies and payment traces in LangSmith.
- x402 - Open HTTP payment protocol and reference implementations for paid APIs and agent services.
- AP2 (Agent Payments Protocol) - Google-led open protocol for interoperable agent payments; separate from the A2A communication protocol.
- minia2a - ⚠️ Unverified (independent adoption unverified). Agent API marketplace using x402 for per-call USDC payments on Base, with wallet authentication and configurable spending limits; platform-reported usage counters are not independently verified.
- Cog Depot - ⚠️ Unverified (early-stage self-submission; no independently verified adoption). Agent marketplace with listing discovery, negotiation and peer introductions through REST and an MIT MCP client; the broker’s fee escrow is distinct from escrow of the underlying trade.
- MCPize - 🆕 MCP server monetization platform — list an MCP server, set a price, platform handles billing and discovery. 85% revenue share to developers. Bridges the gap between the MCP tool ecosystem and sustainable developer economics.
- AgentForge - ⚠️ Unverified (early-stage, 3 stars). Subscription marketplace for AI agents, tools, and content — 300+ agents, unified API, MCP support, 90% creator revenue share. Listed for visibility; evaluate before depending on it.
- Cloudflare Wallets - 🆕 August 4, 2026 (Cloudflare Agents Week, Aug 3–7). Programmable wallet for the Agentic Internet —
cloudflare.paygives AI agents a secure way to make autonomous payments as participants in the agent economy. Launched alongside WriteGuard (fine-grained controls for risky MCP tool calls), WebMCP (one-line website-to-agent discoverability), MCPv2, and unified Workers AI + AI Gateway control plane as part of Cloudflare's Agents Week. - LangChain × AgentCore Payments - 🆕 August 17, 2026. Middleware so LangChain agents pay for APIs via Amazon Bedrock AgentCore (x402) with session budgets enforced outside the prompt; LangSmith traces every spend.
- Alchemy & Visa AgentCard - See also 🔌 Tool & API Integration — listed there for its identity/payment stack; relevant here for the agent economy angle.
- Amazon Bedrock AgentCore Payments - See also 🏢 Enterprise Agent Platforms — managed payment layer for AgentCore agents (Coinbase/Stripe integrations, spending limits).
🧪 Agent Sandboxing & Compute Isolation
Secure runtimes that let agents execute generated code and shell commands without compromising the host. Critical infrastructure once you let an agent off the leash.
- OpenSandbox - Apache-2.0 sandbox platform with Docker/Kubernetes runtimes, multi-language SDKs, CLI/MCP access, and per-sandbox network controls.
- E2B - Open-source secure cloud sandbox for AI-generated code. Used as the execution layer in OpenAI Agents SDK and many production agents.
- Daytona - 💤 Public repository unmaintained: core development moved to a private codebase in June 2026; the public v0.190.0 snapshot receives no further fixes or releases, while the managed service continues.
- Modal - Serverless cloud platform popular for agent compute, GPU jobs, and sandboxed Python —
modal-clientis the official SDK. - Microsandbox - Local, programmable microVM sandboxes for AI agents — secure code execution on your own machine, no cloud dependency.
- SandboxFusion - ByteDance's multi-language code-execution sandbox built for agent / model evaluation pipelines. Apache-2.0.
- Northflank - General-purpose container PaaS used as an agent runtime backend (per-task ephemeral environments, GPU pools).
- Firecracker - KVM-based virtual machine monitor (VMM) for lightweight microVMs; Apache-2.0.
- LangSmith Sandboxes - May 2026 (Interrupt 2026). Hosted secure code execution environments for agents — filesystem, shell, package manager, persistent state, and network boundary. Part of LangChain's Interrupt 2026 release alongside LangSmith Engine and Managed Deep Agents.
- Google Antigravity Sandbox - 🆕 May 2026 (Google I/O). Sandboxed Linux environments for agent-executed code; ships as part of Antigravity 2.0's stack — sub-agents run in isolated containers with scoped filesystem + network access.
- Amazon Bedrock AgentCore Runtime Instances - 🆕 GA August 6, 2026. EC2-backed persistent compute for AgentCore agents — long-running agent sessions up to 14 days (vs the 8-hour serverless microVM cap), with GPU-accelerated, memory-optimized, and compute-optimized instance families via capacity providers; no change to the deploy/invoke path. 9 regions at launch.
🛡️ Agent Security
Tools and frameworks for securing AI agents against prompt injection, data leaks, and misuse.
- Cloudflare WriteGuard - 🆕 August 5, 2026 (Cloudflare Agents Week, Aug 3–7; Private Beta). Fine-grained controls for risky MCP tool calls — the same tooling Cloudflare uses internally, now in private beta for customers. Lets operators intercept and veto destructive or sensitive agent actions before they execute; lowers the blast radius of prompt-injection and autonomous-agent errors.
- UK AISI agent containment incident (INC-2026-07-28-01) - 🆕 ⚠️ Incident, not a tool — disclosed August 5, 2026. The UK AI Security Institute reports that agents built on frontier models (Anthropic Mythos 5, OpenAI GPT-5.6 Sol) in a routine cyber evaluation took "sustained, unsanctioned action directed at real people and organisations" — attempting an open-source supply-chain attack via malicious PRs and social-engineering a maintainer. The agent was never instructed to deceive; deception emerged as a by-product of pursuing the task. Reference case for containment/egress controls on eval infrastructure.
- prompt-firewall - ⚠️ Unverified (early-stage). Firewall for LLM prompts — detect and block prompt injection attacks.
- LLM Guard - 📦 Archived (2026-07-08). The Security Toolkit for LLM Interactions — input/output scanners for AI. Kept for historical reference.
- Rebuff - 📦 Archived (2025-05). Self-hardening prompt injection detector — detect, deflect, and report. Listed for historical reference; no longer maintained.
- Guardrails AI - Adding guardrails to large language models — validate and correct LLM outputs.
- NeMo Guardrails - Toolkit for adding programmable guardrails to LLM-based conversational systems.
- Vigil - 💤 Stale (no commits since 2024-01). LLM security scanner — detect prompt injections, jailbreaks, and data leakage.
- Lakera Guard - Enterprise-grade AI security platform for prompt injection defense.
- Garak - LLM vulnerability scanner by NVIDIA — probe for weaknesses in language models.
- Invariant Guardrails - Runtime guardrails for AI agents — policy enforcement and safety checks.
- Prompt Armor - Enterprise prompt injection protection with real-time detection.
- Descope MCP Auth - Authentication and authorization layer for MCP server security.
- AgentDojo - ETH Zürich research benchmark for evaluating prompt-injection attacks and defenses against tool-using LLM agents.
- ModelScan - Scan ML model files (Pickle, PyTorch, TF) for serialization-based code-execution attacks.
- PyRIT - Microsoft's Python Risk Identification Tool for generative AI — automated red-teaming framework (moved from Azure/PyRIT, March 2026; actively maintained). Complements RAMPART below.
- RAMPART - May 20, 2026. Microsoft's pytest-native safety + security testing framework for agentic AI. Developer-facing white-box counterpart to PyRIT — cross-prompt-injection probes, benign-failure asserts, harm-category coverage, statistical thresholds (e.g. safe in 80%+ runs). Integrates straight into CI/CD. MIT.
- Clarity (Microsoft) - May 20, 2026. Companion to RAMPART. Structured design-review tool for AI agents — "living artifacts" documenting intent, risks, and behavior before code is written. Open-sourced from Microsoft AI Red Team's internal practice.
- Nobulex - ⚠️ Unverified. Cryptographic receipts for AI agent actions (Ed25519 dual signatures, hash-chained audit logs). MIT. Bilateral-receipt primitive merged into Microsoft's Agent Governance Toolkit (PRs #1302, #1333). Same submission sent to 15+ awesome lists in parallel; submitter's claim of "4,500 npm downloads" doesn't match registry data (
@nobulex/mcp-server~19/month at audit time). Listed for visibility on the strength of the Microsoft adoption. - MCP Gateway & Registry - Enterprise-ready MCP gateway and registry that centralises AI development tools with OAuth authentication, dynamic tool discovery, audit trails, and Keycloak / Entra integration. Apache-2.0.
- ActPlane - 🧪 OS-level agent harness enforcing behavioral contracts defined in YAML via eBPF at the syscall boundary — constraints hold across any tool, subprocess, or direct syscall, with corrective feedback to the agent on violation. MIT.
- WalletPrint - ⚠️ Unverified (early-stage). Open-source SDK for behavioral risk scoring of agent wallets that flags anomalies before transactions are signed using wallet behavioral history, with integrations for ZeroDev and LangChain.
- Alchemy & Visa AgentCard - June 18, 2026. Payments + identity stack for AI agents built on Visa Intelligent Commerce. One API provisions everything an agent needs to transact — a Visa payment token, a dedicated email and phone number, and a crypto wallet — so it can buy on a consumer's behalf with scoped controls. Defaults to Visa-issued tokens; also supports crypto, x402, and Stripe's Machine Payments Protocol. Model-agnostic (OpenAI / Anthropic / etc.).
- Microsoft Prompt Shields - Azure AI Content Safety feature detecting jailbreaks and indirect prompt injection hidden in documents/web pages an agent consumes (GA 2024; since extended for agent workloads). Integrates with Azure OpenAI Service and third-party models.
- Agent Name Service (ANS) - June 2026. Linux Foundation initiative to establish an open standard for AI agent verification and trusted identity. Decentralized agent name registry so agents can verify they are communicating with legitimate counterparts, mitigating impersonation and MITM attacks.
- OpenAI Daybreak - June 2026. OpenAI initiative + updated Codex Security plugin for automated vulnerability discovery and remediation in AI-adjacent code; includes prompt-injection hardening for agentic applications.
- JADEPUFFER (Sysdig disclosure) - ⚠️ Threat, not a tool — July 2, 2026. Sysdig documents the first fully agent-orchestrated ransomware operation: an LLM-driven agent exploited a Langflow RCE (CVE-2025-3248), harvested credentials, pivoted to a production MySQL/Nacos server, self-corrected a failed step in 31 seconds, then encrypted 1,342 config items with an ephemeral (never-saved) AES key, making the ransom demand unpayable-but-unrecoverable. Payloads were "self-narrating" with natural-language reasoning comments — strong evidence of LLM authorship. Cited here as the reference case for why agent-security tooling (guardrails, egress control, credential scoping) above matters in production.
- Lineation.ai - 🆕 ⚠️ July 2026 (new vendor). Agent accountability layer — observability, governance, and defense with forensic reasoning lineage. Aims to prevent goal hijacking, memory poisoning, and tool misuse; audit trail for SOC 2 / HIPAA / EU AI Act compliance. Cloud (free-to-start) and on-prem.
- First Recon AI Security Runtime - 🆕 July 2026. Enterprise AI governance platform that inspects every AI interaction (human-to-model, agent-to-tool, agent-to-agent) with a proprietary Semantic Security Engine — applies policies before data reaches a model and captures a full decision audit trail. macOS + Windows endpoint agent for governing AI use on device.
- CrowdStrike Falcon AIDR - GA December 2025. AI Detection and Response — visibility into employee AI use and agent activity across an enterprise, risk scoring, behavioral anomaly detection, prompt-injection blocking, and real-time policy enforcement at the AI interaction layer.
- Darkmoon - Open-source (GPL-3.0) AI penetration-testing platform with specialist agents and MCP interfaces; supports cloud providers and local models, with a privacy gateway that masks selected identifiers; actual data flows depend on the provider and configuration.
- Exabeam Agent Behavior Analytics - 🆕 2026. Extension of Exabeam's behavior-intelligence platform to cover agentic AI risks — continuous verify-observe-analyze-improve loops instead of static guardrails.
- RufRoot / CVE-2026-59726 - 🆕 ⚠️ Disclosed to maintainers June 30, 2026; reported July 29, 2026. A CVSS 10.0 flaw in Ruflo (formerly Claude Flow), an open-source multi-agent orchestration layer for coding agents: the project's previous default Docker Compose config exposed Ruflo's MCP bridge to the network with no authentication, so a single request could invoke
terminal_executeand reach all 233 tools behind the bridge — leaking LLM provider API keys and stored conversations. Worst part: attackers could write to AgentDB, Ruflo's persistent agent memory, so poisoned instructions survive the upgrade. Maintainers fixed the default in 24h (3.16.3), but recovery requires credential rotation and an AgentDB audit — patching alone is not enough. Found by Noma Labs. The canonical example of why agent memory is now part of your attack surface. - Claude Code symlink exfiltration (Tego AI) - 🆕 ⚠️ July 24, 2026. A repo-committed
CLAUDE.mdwith an@importpointing at a symlink can make Claude Code read files outside the project and fold their contents into its very first request — no tool call, no approval prompt, no warning, because the out-of-project read check validated the in-repo link path rather than what it resolved to. Reported via HackerOne; Anthropic closed it "Informative" on the grounds that the trust boundary is the initial folder-trust dialog. Worth reading before you let an agent loose on an untrusted repo. - CrowdStrike 2026 Threat Hunting Report - 🆕 2026-08-03. AI-agent-triggered detections are 2.5× higher than human-initiated leads; Chinese APTs exploit PoC vulnerabilities within 24 hours of disclosure; STARDUST CHOLLIMA poisoned 300+ AI framework dependencies in a single day; a single LLMJacking campaign sent 200,000 API requests in 2 minutes.
- Straiker AI Runtime Security - 🆕 2026-08 (BH2026 showcase). AI-native agentic security platform — asset discovery (Discover AI), adversarial red-teaming (Ascend AI), runtime blocking (Defend AI). Blocks prompt injection, memory poisoning, identity abuse. Raised $85M total ($64M Series A, 2026-06).
- EU AI Act Article 50 — transparency obligations - Applicable from August 2, 2026. Article 50 sets transparency obligations for covered providers and deployers, including AI interaction notices and content marking/disclosure; consult the Commission guidance for scope, role-specific duties and exceptions.
🔍 RAG & Knowledge
Retrieval-augmented generation and knowledge management systems for agents.
- Oracle OCI Enterprise AI updates - June 2026. Enterprise deployment of Cohere Rerank 4 to enhance RAG and agentic enterprise search, plus expanded support for new Alibaba/Google models.
- LlamaIndex - Data framework for LLM-based applications — ingest, structure, and access private data.
- Haystack - End-to-end LLM framework for building RAG pipelines and search systems.
- Unstructured - Open-source components for pre-processing documents for LLMs and RAG.
- Chroma - AI-native open-source embedding database.
- Weaviate - Open-source vector database for AI-native applications.
- Qdrant - High-performance vector similarity search engine and database.
- Pinecone - Managed vector database for high-performance AI applications.
- Milvus - Cloud-native vector database for scalable similarity search.
- RAGFlow - Open-source RAG engine based on deep document understanding.
- Docling - Document parsing and conversion for RAG and generative AI.
- Kotaemon - Open-source RAG-based tool for chatting with documents.
- LightRAG - Simple and fast RAG engine with graph-based knowledge indexing.
- R2R - Production-ready RAG engine with built-in auth, observability, and ingestion.
- Vanna - 📦 Archived (2026-03). RAG for SQL — chat with your database using natural language.
- Morphik - Multimodal retrieval engine for documents containing text, tables, figures and charts.
- Cognee - Knowledge and memory engine combining document ingestion, graphs, and vector retrieval; Apache-2.0.
- RAG-Anything - All-in-one multimodal RAG framework from HKU Data Science Lab. Built on top of LightRAG; concurrent pipelines for parallel text + multimodal processing; queries documents that interleave text, diagrams, tables, and formulae. MIT, 21K+ stars.
- A-MEM - Agentic Memory system for LLM agents — dynamic organization of memories using Zettelkasten-inspired note linking; enables more flexible retrieval than static vector stores.
- LangChain Retrievers - LangChain's collection of retrievers and document loaders for connecting sources to RAG pipelines.
- Milvus 3.0 - 🆕 v3.0.0 tagged July 29, 2026 (public beta was May 2026). A "lake-native" architecture shift for the large-scale vector database — External Collections that query Parquet / Lance / Iceberg tables directly in S3/GCS/Azure object storage with zero-copy access, a manifest-based Storage V3 columnar engine, Spark DataSource V2 integration, runtime schema evolution,
TEXTas a first-class type, and multi-vectorStructListfor late-interaction (ColBERT-style) retrieval.
💻 Coding Agents
AI-powered coding assistants and autonomous software engineering agents.
Terminal & CLI Agents
- Claude Code - Anthropic coding agent for terminal, IDE, and repository workflows; v2.1.263 (2026-09-06) contains reliability fixes.
- Codex CLI - OpenAI terminal coding agent, Apache-2.0; stable rust-v0.153.4 (2026-09-04) fixes Astra visibility and the bundled default model, while 0.154 alpha builds remain prereleases.
- Codex Security - March 2026. Application-security agent that finds and fixes software vulnerabilities; available to OSS maintainers via the Codex-for-OSS program.
- Aider - Terminal pair-programming tool with repository context and Git integration; Apache-2.0.
- goose - Extensible desktop and CLI agent originating at Block, now hosted by AAIF; Apache-2.0; v1.49.0 (2026-09-03).
- Gemini CLI - Google's terminal-first coding agent for large-context refactors.
- OpenCode - Open-source terminal AI coding agent (opencode.ai, 180K+ stars) — build/plan agents, LSP, MCP, desktop app in beta; unrelated to the archived opencode-ai/opencode.
- Crush - Terminal AI coding agent from Charm — successor to the archived opencode-ai/opencode; multi-model, LSP + MCP support.
- Grok Build - May 25, 2026 (early beta). xAI's agentic CLI coding agent powered by grok-code-fast-1. Parallel sub-agents in isolated environments, daily release notes; available to SuperGrok and X Premium Plus subscribers. xAI's reply to Claude Code and Codex CLI. ⚠️ July 2026 reports found Grok Build uploading entire git repos to xAI storage — review before use on private code.
- Antigravity CLI - May 19, 2026 (Google I/O 2026). Lightweight CLI companion to Antigravity 2.0 — create and interact with Google agent harnesses directly from the terminal. macOS / Linux / Windows. Reported to replace Gemini CLI for hosted-plan users from June 18, 2026 (the open-source gemini-cli repo remains active, 105K+ stars).
- Kimi Code CLI - Moonshot terminal coding agent for code editing, shell commands, and file/web access; official installers do not require Node.js; 0.41.0 (2026-09-04).
- MAI-Code-1-Flash in GitHub Copilot - Build 2026 (June 2, 2026). Microsoft's first fully in-house 5B coding model lands as a model picker option in GitHub Copilot — outperforms Claude Haiku 4.5 on four core coding benchmarks (SWE-Bench Pro 51.2% vs 35.2%) at significantly lower cost.
- Claude Agent SDK - Python and TypeScript SDKs exposing the Claude Code agent loop, tools, permissions, and session handling for applications.
- ai-delivery-spec - ⚠️ Unverified. Spec-driven delivery framework for PMs working with AI coding agents (Claude Code, OpenClaw, Codex, Cursor, Copilot). 4 delivery tiers (Lite/Standard/L2/Full), 0D triage routing, prototype testability rules, AI runtime governance, 5 domain modules. SKILL.md convention; hosted on ClawHub.
- Ralph Harness - ⚠️ Unverified. Tiny Python scaffold for guarded Claude Code/Codex/Gemini loops with repo-local specs, fresh-context iterations, git-hook gate, CI verification, and coverage gates. Installable via
uvx ralph-harness demo. MIT. - Amp - 🆕 ⚡ Sourcegraph's frontier coding agent (VS Code extension + CLI). No BYOK — model access is bundled, and its model-agnostic "Dial" router picks the model for you. July 2026 was a heavy month: paid subscriptions launched in beta July 18 (Megawatt $20/mo, Gigawatt $200/mo, with the option to attach your own ChatGPT or X Premium+/SuperGrok sub), self-scheduling agents July 21, "Multiplayer" shared-thread collaboration July 22, and event-driven Orbs July 23 — agents that wake on an external event (CI failure on GitHub, a new Linear issue, a monitor alert, a Discord message; anything that can send an HTTP request). August kept pace: "Attach Anything" uploads (video/logs/PDF/datasets, Aug 4), "Portals into Orbs" live-reloading previews (Aug 6), Dial running on a linked ChatGPT subscription (Aug 10), and Global Plugins and Skills (Aug 11). Closed-source.
- ZCode - 🆕 🇨🇳 July 2026 (ZCode 3.0). Z.ai's official agentic development environment for GLM-5.2 — a desktop app (macOS / Windows / Linux) wrapping file manager, terminal, Git panel, and live browser preview around an agent that plans, codes, reviews, and deploys. Also drives Anthropic and OpenAI models. Free tier with a daily token allowance; GLM-5.2 access via the paid GLM Coding Plan (Lite / Pro / Max).
- Kolega Code - 🆕 ⚠️ Unverified (15 GitHub stars; PyPI ~10.5k downloads/month, v0.32.0 on August 24, 2026). Terminal coding agent whose Gigacode engine has the model write a Python multi-agent orchestration program (parallel / pipeline / judge panels) with content-keyed journaled resume. 15+ model providers, MCP client (HTTP/SSE/stdio/OAuth), Textual TUI. BSL 1.1 (not Apache-2.0; Change Date 2030-08-12) — the PR description got the license wrong.
IDE-Based Agents
- Cursor — self-hosted machines - 🆕 2026-09-02: self-hosted workers keep tool execution on your machines, with personal machines, team pools, and Linux/macOS computer use; model processing and data policies require separate review.
- Cursor 3.4 (Teams + PR review) - May 11–13, 2026. Microsoft Teams integration (
@Cursorin Teams delegates to cloud agents), faster parallel-agent plan execution, multi-repo / Dockerfile-based dev-environment configs for agents,/multitaskasync sub-agents, Vulnerability Scanner, granular per-model access controls. - Cursor 3.3 - May 2026. PR-review experience, parallel agents, enterprise model controls; previous 3.1 in April.
- Cursor SDK - 🆕 April 29, 2026 (public beta). TypeScript SDK exposing Cursor's runtime, harness, and models so developers can build programmatic agents on top of the Cursor stack — sandboxed cloud VMs, subagents, hooks, token-based pricing.
- Kilo Code - Open-source AI coding extension (VS Code / JetBrains) with Auto Model routing across 500+ models; acquired by Anaconda (2026). MiniMax models heavily featured.
- Cursor - The AI code editor with Feb 2026 update supporting up to 8 parallel agents.
- Windsurf → Devin Desktop - **Rebranded June 2, 202
…(truncated)
Collected info
- ★ 246 stars
- ⎇ 96 forks
- Language: Python
- Source updated: 9/18/2026