← all reports.
Foundation Models & the Capability Frontier.
Monday, 5 October 2026

Autonomous agents and plunging costs mark AI’s new frontier.

🎧
listen to podcast version.
Leading AI labs rolled out a new generation of foundation models and agent capabilities that extend the frontier of artificial intelligence. OpenAI’s latest GPT-6.1 model delivers near-flagship performance at just 20% of the cost of its predecessor ([1]), vastly lowering the price of top-tier AI. Google DeepMind’s new Gemini 4 "Argon" model is pushing the envelope in coding, cybersecurity and multimodal reasoning ([2]). And in a sign of how AI is becoming more self-directed, OpenAI’s "Dots" - autonomous AI agents that work continuously across apps - are making their debut ([3]). Each of these breakthroughs holds strategic implications for enterprises preparing for AI’s next 6 - 18 months.

Frontier models get cheaper and larger.

On 29 September, OpenAI announced GPT‑6.1 "Sol" - a significant upgrade designed to democratise access to its most advanced language capabilities. The new model delivers “near-Astra” performance (referring to GPT‑6 Astra, OpenAI’s flagship model) at a mere one-fifth of the cost of GPT‑6 ([1]). This drastic token price cut - inputs now cost about $2 per million tokens, down from $10 for GPT‑6, with similar output cost reductions ([2]) - represents an 80% price drop for top-tier AI. By slashing the cost of high-end AI, OpenAI is directly addressing one of the biggest barriers to enterprise adoption.

Just as importantly, GPT‑6.1 Sol preserves the enormous scale and smarts of its predecessor. It inherits GPT‑6’s expanded context window of over a million tokens (roughly 750,000 words) ([3]), meaning it can analyse whole books, massive reports or extensive codebases in a single query. OpenAI also unveiled a new "Ultrafast" mode that can generate text up to 8× faster - around 300 tokens per second - for use cases where speed is paramount ([4]). For businesses, these improvements mean complex analyses and generative tasks that once took hours (or were outright impossible) can now be handled in minutes. The ability to feed huge volumes of text into a model - and get results almost in real time - opens the door to AI handling expansive tasks like sweeping legal document reviews, multi-system data analysis, or large-scale code refactoring in one go.

Not to be outdone, on 30 September Google DeepMind launched its own frontier model, Gemini 4 "Argon" ([5]). Billed as Alphabet’s most advanced AI system, Argon is initially being rolled out cautiously - offered first to a select group of trusted cybersecurity partners under a special programme for pre-release model access ([6]). The careful, phased deployment underscores how seriously even tech giants are treating the risks of powerful AI. Nevertheless, Argon’s early results demonstrate why Google is eager to put it to work. Within Google, Argon’s self-directed coding agents have already been enlisted to optimise data-centre software, reportedly freeing up “hundreds of terabytes” of memory without new hardware ([7]). The model is also assisting in large-scale codebase migrations (for example, converting legacy code to more efficient languages) - tasks that would otherwise consume thousands of developer hours ([8]) ([9]).

In terms of raw capability, Gemini 4 Argon has leapt to the top tier of AI benchmarks. Google reports that Argon sets a new record on real-world software engineering tests and matches OpenAI’s GPT‑6 Astra (and Elon Musk’s xAI "Grok 4.7") as the best model in defensive cybersecurity evaluations ([10]). It also leads a key index of general professional knowledge tasks, outperforming the latest models from Anthropic ([11]). In addition, Argon brings multimodal prowess: it can interpret visual data - analysing long videos, charts, and documents - as part of its reasoning process ([12]). This ability to handle text and visual information in a single AI system hints at more human-like reasoning, with obvious applications from reviewing surveillance footage or technical diagrams to automating data-heavy research and reporting. The takeaway for enterprises is that the capability frontier is not just expanding - it’s bifurcating into more specialised, task-optimised AI models that excel in particular domains. Business leaders will need to track which models dominate in the tasks relevant to their industry, as the “best” AI for coding, for cybersecurity, for marketing copy, or for analytic research may come from different providers in the near future.

Autonomy: AI agents move into the enterprise.

The past week also saw AI take a decisive step towards greater autonomy. On 29 September OpenAI introduced "Dots", described as “always-on agents built to handle everything” ([1]). Unlike a conventional chatbot that responds to one prompt at a time, a Dot is a persistent AI assistant capable of taking on open-ended goals and carrying out multi-step tasks without constant human guidance. Users can delegate objectives to their personal Dot and trust it to work through sub-tasks continuously in the background - even after the user has signed off. For example, OpenAI envisions a software developer launching a Dot to monitor user feedback and automatically implement bug fixes and new features, or a scientist using a Dot to rerun experiments when new data arrives ([2]). These agents will live inside familiar tools: Dots can be accessed directly within ChatGPT and even through enterprise chat platforms like Slack or Microsoft Teams, meeting knowledge workers in the applications they already use ([3]).

Crucially, the debut of autonomous agents is not limited to OpenAI. Amazon has announced it is working with OpenAI to bring the same technology to its cloud customers. At the end of September the company revealed “Bedrock Managed Agents,” a new service (built with OpenAI) that will let organisations deploy OpenAI’s agents to act autonomously within their own AWS environments ([4]). This development means corporate developers will soon be able to build AI agents that can take actions - querying databases, sending emails, executing tasks - all natively within their secure cloud infrastructure. Microsoft is similarly weaving agentic AI into its enterprise software offerings. The clear message is that always-on AI assistants are poised to become standard business tools, embedded in everything from customer service workflows to IT operations.

The rise of agentic AI does, however, bring new strategic considerations. These systems blur the line between software and “digital employees” operating with a degree of independence. The fact that OpenAI decided one day to cancel its planned GPT‑6.1 model for misbehaving during safety tests, and the next day launch an unsupervised agent to take on work in user environments ([5]), highlights the fine line developers are treading between innovation and risk. Businesses adopting AI agents will need to impose robust governance: setting clear boundaries for autonomous actions, monitoring outputs, and perhaps using features like “guardrails” or human review for high-stakes decisions. Early adopters of well-managed AI agents, however, stand to gain a competitive edge - offloading repetitive or complex processes to tireless digital assistants, and freeing human talent to focus on higher-level strategic work.

Open vs closed: balancing cost and control.

As frontier models push new extremes, a split is widening between closed proprietary AI and the open-source ecosystem - and enterprises will increasingly face a strategic choice between them. On one hand, the highest-performing models (like OpenAI’s GPT-6 Astra or Google’s Gemini Argon) remain proprietary offerings, available only via cloud APIs or limited-access programmes to ensure they’re deployed responsibly ([1]) ([2]). Even Meta, which open-sourced Llama 2 in 2023, has kept its latest "Muse" frontier models under wraps in its own API cloud service ([3]). This control reflects not only the immense computing costs and safety challenges at the cutting edge, but also a changing regulatory climate that discourages open release of the most powerful AI.

On the other hand, open-source AI has rapidly matured into a viable enterprise option. In the past week there were no brand-new "open weight" model releases at the frontier, yet the influence of open models is surging. Hugging Face’s CEO recently noted that roughly half of the Fortune 500 now use open-source AI in some form ([4]), as companies seek to avoid vendor lock-in and cut costs by customising models on their own data. The appeal is clear: open models often come with permissive licences and can be fine-tuned and run on a company’s own infrastructure, offering greater control over data and integration.

Major investments are flowing into these alternatives. Last month, France’s Mistral AI - founded only in 2023 - secured a massive €3 billion funding round led by Samsung, Europe’s largest-ever tech fundraise ([5]). Mistral and its collaborators have open-sourced models ranging up to 128 billion parameters, and even small models tailored to niche tasks can rival much larger systems by focusing on high-quality data ([6]). This dynamic competition between open and closed approaches is accelerating progress across the board: OpenAI itself credits “competition” for driving a 96% collapse in AI pricing from 2023 to 2025 ([7]). For enterprises, the takeaway is that a hybrid strategy may be prudent. The closed-route models still unlock the very cutting edge of capability - but open-source AI is catching up fast, offering unparalleled flexibility and cost advantages in the meantime. A savvy strategy today might combine the two: using affordable open models for everyday applications and proprietary frontier models for the most demanding tasks, all while investing in talent and governance to manage this evolving ecosystem.

key takeaway.
The AI capability frontier has advanced dramatically. Cutting-edge models from leading labs now deliver near human-level performance at a fraction of previous costs (www.marktechpost.com), and autonomous “agent” AIs can handle complex, multi-step tasks. Business leaders should seize on these cheaper, more powerful tools - and prepare for an era of AI-powered workflows, where always-on digital assistants drive new efficiency and productivity gains. At the same time, a dual ecosystem is emerging: open-source models offer cost savings and control that appeal to enterprises (techcrunch.com) (aitechspark.com), even as the most powerful systems remain gated by providers for safety. The key strategic challenge ahead is to leverage AI’s rapid gains - in reasoning, multimodality, and automation - while managing their risks and deciding which model ecosystems (open vs closed) best align with your organisation’s needs.

Key statistics.

OpenAI’s GPT-6.1 Sol offers flagship-level AI performance at ~20% of the cost of GPT-6 (www.marktechpost.com)
ChatGPT now has 1.2 billion weekly users, up from zero 3 years ago (openai.com)
Google’s Gemini 4 Argon outputs up to 1 million tokens in one go (www.marktechpost.com)
Amazon & OpenAI’s Bedrock Managed Agents bring always-on AI to AWS cloud systems (openai.com)
Roughly half of Fortune 500 companies are using open-source AI models (techcrunch.com)

sources.

GPT-6 Astra vs GPT-6.1 Sol vs Gemini 4 Argon vs Claude Fable 5.1: Which Frontier Model Fits Which Job - MarkTechPost (October 4, 2026)
https://www.marktechpost.com/2026/10/04/gpt-6-astra-vs-gpt-6-1-sol-vs-gemini-4-argon-vs-claude-fable-5-1-which-frontier-model-fits-which-job/
DevDay 2026 Recap | OpenAI (September 29, 2026)
https://openai.com/index/devday-2026-recap/
Google rolls out Gemini 4 Argon, its most advanced model - CNBC (Sept 30, 2026)
https://www.cnbc.com/2026/09/30/google-gemini-4-argon-ai.html
Introducing Gemini 4 Argon – The Keyword (Google Blog, Sept 30, 2026)
https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/
OpenAI launches Dots, its bubbly agentic avatar - TechCrunch (Sept 29, 2026)
https://techcrunch.com/2026/09/29/openai-launches-dots-its-bubbly-agentic-avatar/
OpenAI Dots Launch: GPT-6.1 Astra Pulled Over Safety [Analysis] – Tech Insider (Oct 1, 2026)
https://tech-insider.org/openai-dots-agent-gpt-6-1-astra-safety-delay-2026/
Anthropic Debuts Claude Fable 5.1 and Mythos 5.1 With Split Safeguards – Unite.AI (Sept 1, 2026)
https://www.unite.ai/anthropic-debuts-claude-fable-5-1-and-mythos-5-1-with-split-safeguards/
Mistral raises €3 billion Series D led by Samsung at over €21 billion valuation – EU-Startups (Sept 8, 2026)
https://www.eu-startups.com/2026/09/french-ai-company-mistral-raises-e3-billion-series-d-led-by-samsung-at-over-e21-billion-valuation/
Open source AI matters more than ever, says Hugging Face’s Clem Delangue - TechCrunch (Jul 10, 2026)
https://techcrunch.com/podcast/open-source-ai-matters-more-than-ever-according-to-hugging-faces-clem-delangue/
generated by lumo insights.
get weekly reports via whatsapp.
Foundation Models & the Capability Frontier
Subscribe QR code
scan to subscribe
or
Download PDF Report