AI Developments — Thursday, September 24, 2026
Anthropic Claude Code
This Week in Anthropic Claude Code
Claude Code is Anthropic's AI that lives right in your terminal, reading your files, running commands and writing real code instead of just chatting about how it would do it. This week's news split neatly into two stories: a smarter, cheaper brain running underneath, and a serious rethink of how Claude Code handles big jobs with more than one moving part.
On September 22nd, Claude Opus 5.5 became the new default Opus model across Claude Code, the Claude app and Cowork for Pro, Max and Team users, and it's a genuine upgrade dressed up as a cost cut: a full million-token context window, priced at $4 per million tokens in and out with cache reads at a fraction of that, while landing performance close to Anthropic's newest "Fable"-class model for roughly 40% less than the older Opus 5 cost to run. Five days earlier, on the 17th, Anthropic opened a beta redesign of Claude Code Projects that changes what a "project" actually means: instead of one linear conversation, a single project can now spin up several parallel threads, each one a full cloud session running on its own git branch and its own copy of the repo, all pulling from the same shared memory — CLAUDE.md files, AGENTS.md files and notes Claude has generated along the way — so a coordinator conversation can hand out pieces of a big job to multiple workers running on Opus at high effort and have them report back when they finish, instead of one thread grinding through everything in sequence.
Smaller but very real fixes rounded out the week too. Version 2.1.275 added a "send-now" keyboard shortcut for interrupting a turn mid-thought and started syncing skills and plugins straight from a claude.ai account into terminal sessions, and a quick patch the next day, 2.1.276, fixed a bug that made every single request fail outright for teams routing Claude Code through their own proxy or gateway server — the kind of silent breakage that can quietly take down an entire company's workflow until someone notices.
None of this is a single "look what it can do now" demo moment — it's Anthropic making the same tool cheaper to run at its best while making it better at juggling several jobs at once without a person babysitting every thread. The thing worth watching next is whether that Projects beta, still limited to a slice of Pro and Max users, expands fast enough to become the normal way people use Claude Code, or stays a power-user feature tucked behind a waitlist.
Xai Grok
This Week in Xai Grok
Grok is Elon Musk's chatbot, wired directly into X and known for a looser personality than most AI assistants. This week the story that had been teased and then delayed for weeks finally landed: the model itself.
Grok 4.7 launched on September 21st, landing on the xAI API, inside Grok Build and through Cursor, built around a bigger base model, a longer reinforcement-learning training run, and a fix for a genuine bug Musk admitted to in public. During training, penalties for long answers had accidentally taught the earlier version of the model to bail out of hard problems early and skip double-checking its own work, a bit like a student erasing a correct answer because it took too long to reach it; the extra weeks of training that caused the delay were spent specifically teaching Grok to keep working on hard problems and check its own answers more carefully before handing them over. It ships with a 500,000-token context window, reads both text and images even though it only writes text back out, and it's priced the same as its predecessor at $2 per million input tokens and $6 per million output tokens, so a release built around "we fixed the model to actually try harder" didn't come with a price hike attached. Musk said afterward on X that Grok 4.7 has kept climbing in independent rankings since launch, evidence the extra weeks of tuning paid off. Alongside the flagship, xAI also quietly shipped Grok Voice Transcribe 2.0, a speech-to-text model built for stronger real-world accuracy across multiple languages with word-level timestamps, while Grok Build, the company's coding-agent tool, kept picking up steady polish in the background.
With Grok 4.7 finally out the door, attention is already sliding further down the road: Musk says Grok 4.8 is finishing training this week and Grok 4.9 should catch up to competing models from OpenAI and Anthropic, but he's saving the "artificial general intelligence" label specifically for Grok 5, a rumored 6-trillion-parameter model still without a firm release date. Whether that next leap actually arrives on schedule, after 4.7 already slipped by more than a week, is the thing worth watching next.
AI Developments — Wednesday, September 23, 2026
Anthropic Claude CoWork
This Week in Anthropic Claude CoWork
Claude Cowork built its reputation on being the Anthropic tool you could hand a big, messy job to and then walk away from — cleaning up files, digging through research or drafting a full report while you went and did something else, running on the same agent engine as Claude Code but without any terminal involved. This week Anthropic made the biggest change to Cowork since it launched: it stopped being a separate thing altogether.
On September 16th, Anthropic announced it was folding Cowork straight into ordinary Claude chat, cutting what used to be three separate modes — chat, Cowork and Claude Code — down to two. The idea is that nobody should have to guess upfront whether their question is a five-second answer or a twenty-minute research project; you just type what you need, and Claude figures out under the hood whether to reply right away or quietly go off and do the bigger job, all while holding onto the same context, skills and connectors you already had set up. Riding along with the merger were two brand-new tools, Claude Docs and Claude Slides, both in beta, which let you write a document together with Claude paragraph by paragraph or ask for a presentation and get a full slide deck drafted back, downloadable as an actual PowerPoint file or a PDF. Claude Design, the tool for building websites and visual prototypes that Anthropic shipped back in April, now works straight inside any conversation too, instead of living off in its own separate corner of the app.
The rollout itself is deliberately slow and staged: Pro and Max subscribers are getting the unified experience first, over the coming weeks, with Team and Free plans following after that, and Enterprise administrators getting a mandatory 30-day heads-up before anything changes for their organizations, since a company that's built workflows around the old three-mode setup needs time to adjust. The timing is a little ironic, too, because just over a week earlier a Windows 11 security update had broken Cowork's ability to reach local files entirely, cutting off the shared-drive connection it needs to run commands on your own machine, until Microsoft shipped a fix on September 14th — a reminder of just how much plumbing sits underneath a feature that's supposed to feel effortless.
None of this is a flashy new model, but it's a real bet from Anthropic that the fiddly decision of "which mode do I need" was actually getting in people's way. The thing worth watching next is whether folding everything into one Claude makes the product feel simpler the way Anthropic hopes, or whether power users who liked having Cowork as its own dedicated space for long, unattended jobs end up missing the separation.
Eleven Labs
This Week in Eleven Labs
ElevenLabs made its name cloning voices convincingly enough to fool your own grandmother, and it's since branched out into dubbing video, transcription and, more recently, AI-generated music. The past seven days were unusually calm for a company that normally ships something new every few days, so this week's story is really about the last genuinely big thing ElevenLabs did, which is still sending ripples through the music industry.
Quiet week for Eleven Labs — no major announcements in the past seven days. The last notable move was a one-two punch in mid-September: on September 10th, ElevenLabs signed a multi-year licensing deal with Universal Music Group to build a separate AI music platform for fans, built on licensed songs and actual artist participation, and the very next day it launched Music v2.5, which it's calling its best music model yet, with more realistic-sounding instruments and songs that hold together better from start to finish. ElevenLabs ran a blind test with nearly 48,000 side-by-side pairs and says listeners preferred the new model's takes most of the time, and the company went out of its way to point out that no UMG artist's music was used to train it, keeping the licensing deal and the model launch legally and technically separate even though they landed a day apart.
Worth watching next: whether that separation between the UMG deal and the actual model holds up as the fan-facing music platform gets built out, and whether other major labels follow Universal's lead now that ElevenLabs has shown it can ship a chart-quality model without leaning on anyone's back catalog at all.
🔍 Tool Spotlight
Sakana AI Fugu Max
Most AI headlines these days are about a single model getting even bigger, but Sakana AI, a startup founded by former Google engineers, took the opposite approach with Fugu Max, released alongside its bigger sibling Fugu Ultra v2 on September 11th. Instead of building one enormous brain, Fugu Max is an orchestrator: a lightweight coordinator called TRINITY, weighing in at a tiny 0.6 billion parameters, that reads your request and hands pieces of it out to a pool of other AI models, giving each one a role as a Thinker, a Worker or a Verifier, then stitches their answers back together into a single response. Think of it less like hiring one genius employee and more like a sharp project manager who knows exactly which specialist on the team should handle which part of a job, then double-checks everyone's work before it goes out the door. The payoff is real money: Sakana says Fugu Max beats rivals like Sonnet 5 and Kimi K3 on several benchmarks while costing 40 to 60 percent less, at just $2 per million input tokens. It's aimed squarely at developers and companies burning through API budgets on frontier models, and the surprising part is how little raw horsepower the coordinator itself actually needs to pull the whole trick off.
Source: MarkTechPost / Sakana AI, September 2026
AI Developments — Tuesday, September 22, 2026
Anthropic Claude
This Week in Anthropic Claude
Claude is Anthropic's AI assistant, the one people lean on to draft emails, debug code and dig through mountains of paperwork, and increasingly the one companies plug straight into their own software instead of typing into by hand. The strangest headline of the week landed on September 17th, when Anthropic revealed that Claude is now doing a meaningful chunk of the work of building its own successor, handling about 26 percent of the model research and development that goes into the next version of itself. That doesn't mean the AI slipped its leash and went rogue in a basement server room; Claude is completing whole research tasks end-to-end from a single high-level instruction, but a human is still checking the work at every step, more like a very fast, very tireless research assistant than an independent scientist working alone.
The rest of the week's news was less philosophical and a lot more practical. Claude rolled out a big batch of new Small Business tools, adding 43 ready-made workflows and 27 connectors into everyday software like Shopify, Salesforce, Zoom, Xero, Gusto, Square, Stripe and Zapier, so a small shop owner can have Claude draft a sales proposal, chase a lead or put together a report without hiring a developer to wire everything together first. On the more technical side, the Claude Developer Platform picked up budget controls and geo-pinned inference, meaning a company can now cap what an AI agent is allowed to spend and choose which part of the world its requests actually run in, plus new settings that let developers wall off exactly which websites a Claude agent's search and browsing tools are allowed to visit, turning "the AI can browse the internet" into "the AI can browse this specific, approved slice of it." The Messages API also gained the ability to compact a long conversation on demand, essentially letting a chat trim its own history so it doesn't get slower or pricier the longer it runs.
Underneath all of that, Anthropic quietly graduated its Admin API's user-management tools out of beta for Enterprise customers, the unglamorous kind of update that lets a company's IT department actually manage who has access to what without crossing their fingers that a beta feature won't break on them.
None of this is a single splashy new model, but together it's a picture of Anthropic pushing Claude in two directions at once: further into the guts of its own next model, and further into the unglamorous back office of ordinary small businesses. The thing worth watching next is how comfortable people get with the idea of an AI helping design the AI that replaces it, and whether that 26 percent figure keeps climbing.
Perplexity AI
This Week in Perplexity AI
Perplexity is the AI search engine built around a simple pitch: ask it a question and it answers in plain English with real sources cited underneath, instead of handing you ten blue links and leaving you to do the digging yourself. This week Perplexity spent less time on search itself and more time getting its tools installed on the actual computers people already use every day.
The biggest move came through a new partnership with HP, announced September 15th, that has HP pre-loading the Perplexity Windows app straight onto new PCs, so instead of downloading anything, sourced, cited answers and Perplexity's connected apps are just sitting there waiting the moment someone unboxes a new laptop. That rollout builds on Portable Computer for Windows, a feature that lets Perplexity run recurring tasks and longer workflows locally on a person's own machine rather than off in the cloud somewhere, which matters for two reasons: sensitive work stays on the device instead of getting shipped out over the internet, and those local tasks don't chew through a user's paid credits the way cloud-run ones do.
Perplexity Computer, the company's broader push into actually doing work rather than just answering questions, also picked up effort controls on September 17th, a simple dial running from Light to Ultra that matches how hard the underlying model thinks, and how much it costs, to how complicated the task actually is, so someone summarizing a short email isn't burning the same computing power as someone researching a twenty-page report. That same day, Perplexity published a practical guide called "AI in the Workplace," walking through ten common tasks knowledge workers deal with every week, complete with ready-to-use prompts and safety tips, clearly aimed at people who've heard of Perplexity but aren't quite sure what to actually type into it.
Taken together, it's a week about meeting people where they already are, on a Windows laptop, at a desk job, without much AI experience, rather than chasing the flashiest new search trick. The thing worth watching next is whether the HP deal is the first of more hardware partnerships, the kind of quiet distribution move that ends up mattering more in the long run than any single feature launch.
AI Developments — Monday, September 21, 2026
Google Gemini
This Week in Google Gemini
Gemini is Google's AI assistant, built right into Search, Gmail, Android and the Chrome address bar, which means when it changes, that change shows up inside tools you already use every day instead of some separate app you have to remember exists. This week's moves weren't about one big splashy model launch; they were about giving Gemini better ears, a tidier memory for organizing your stuff, and a sharper toolkit for the developers building on top of it.
The most visible update lands in ordinary Gemini apps: a new Notebooks feature started rolling out on September 14th, letting you gather sources into named notebooks with their own focus area, up to ten sources apiece, so instead of scrolling back through an endless chat history to find that one article you asked about last week, you can just reopen the notebook you built for it. Google also shipped Gemini 3.8 Live, a pair of audio-to-audio models built specifically for real-time voice apps, meaning a voice assistant powered by them can listen and respond instantly without the awkward pause of translating speech to text and back again first, which matters if you're trying to have an actual back-and-forth conversation instead of a call-and-response. On the developer side, Google rolled out an update to its Antigravity coding-agent tooling that changes how its tools talk to each other under the hood, and it released two new open Gemma models, a 26-billion and a 31-billion parameter version, giving developers a smaller, freely downloadable alternative to the full Gemini models for projects that don't need the biggest brain in the building.
None of this is the kind of update that trends on its own, but stack it up and Gemini is quietly getting better at three unglamorous things: keeping your research organized, talking back instantly, and giving developers more size options to build with. The thing worth watching next is whether that Notebooks feature actually changes how people use Gemini day to day, or ends up as one more organizational tool that sits unused.
Meta Llama
This Week in Meta Llama
Llama is Meta's family of AI models, well known for being free to download and build with instead of locked behind a paid app, which made it the default starting point for years for developers who wanted a genuinely capable AI without a subscription attached.
Quiet week for Meta Llama — no major announcements in the past seven days, and stretching the search back two weeks barely turned up more. The last real news affecting Llama's future wasn't even a new Llama release: back in April, Meta Superintelligence Labs shipped Muse Spark, a closed, API-only model built to power Meta's own chatbots instead of Llama, with more proprietary Muse models following since as part of a broader reorganization that has reportedly pushed back whatever Meta has planned next for Llama itself.
The bigger question worth watching isn't a feature drop, it's whether Meta ever ships another flagship open-weight model under the Llama name at all, or whether Llama quietly settles into being the smaller, older sibling while all of Meta's newest muscle goes into models kept behind closed doors.
AI Developments — Sunday, September 20, 2026
Microsoft Copilot
This Week in Microsoft Copilot
Microsoft Copilot is the AI assistant Microsoft has been sprinkling into basically everything — Word, Excel, PowerPoint, Outlook, Teams, even the file explorer — so that instead of switching apps to ask a chatbot for help, the help just shows up wherever you're already working. This week's updates were part of a broader push Microsoft calls "Wave 3," and the theme was giving people more say over which AI brain actually answers their question.
The most noticeable change is a redesigned model picker: it used to live tucked in the top-left corner, and now it sits right inside the box where you type your prompt, split cleanly into a GPT section and a Claude section, with whichever models your school or workplace has turned on. That pairs with an upgrade to Researcher, Copilot's deep-research tool, which now offers a "Council" mode that runs your question through both GPT and Claude at the same time, keeps each one's full answer, and then adds a summary pointing out where the two agree, where they disagree, and what each one caught that the other missed — like getting two experts to double-check each other's homework instead of trusting just one. Copilot in Excel also got sharper for anyone doing serious number-crunching, pulling in live market data and financial research directly into a spreadsheet and adding clearer tracking of exactly what changed and why, which matters a lot when the numbers in question are somebody's actual budget or investment plan. Meanwhile Copilot Cowork, the tool that runs longer independent tasks in the background, picked up cost visibility so people can actually see what a task is costing before it eats up a budget, plus a rebuilt Notebooks interface and the ability to share a useful "skill" you built with teammates instead of reinventing it every time. On the more everyday side, Copilot is rolling out the ability to create and check Planner tasks straight from a chat, and IT admins are getting new exportable usage reports so they can actually see how Copilot is being used across an entire organization.
None of this is a single splashy headline moment — it's Microsoft quietly making Copilot feel less like one fixed assistant and more like a control panel where you pick your AI, watch what it costs, and hand off real chunks of work with some actual visibility into what happened.
Artlist
This Week in Artlist
Artlist started out as a place to license music and video clips for free of the copyright headaches that usually come with putting a song under your YouTube video, and over the past couple of years it's grown into a full AI creative toolkit, letting people generate images, video, voiceovers and music instead of just licensing stuff other people made.
Quiet week for Artlist — no major announcements in the past seven days, and broadening the search out to two weeks didn't turn up much more, either. The last genuinely notable move was back in June, when Artlist announced it was cutting roughly 200 of its 500 employees, about 40% of the company, as part of a shift toward what it called an "AI-native" way of running the business, with CEO Ira Belsky pointing out that competitors were running lean AI-first teams a fraction of Artlist's size. Before that, the bigger story was the April launch of Artlist Studio, a tool built to give creators shot-by-shot directorial control — casting, camera angles, continuity — over AI-generated video instead of just typing a prompt and hoping for the best.
Worth keeping an eye on: whether that AI-native restructuring actually translates into new features landing faster, now that the company has explicitly bet its whole organization on moving quicker with a smaller team.
AI Developments — Saturday, September 19, 2026
SAPJoule
This Week in SAPJoule
Joule is SAP's AI assistant, and if you've never heard of it that's because it doesn't live on your phone — it lives inside the giant business software that runs the finance, supply chain and HR departments of some of the biggest companies on Earth. Think of it less like a chatbot you type fun questions into and more like a very capable coworker who sits inside a spreadsheet the size of a small country's tax records, and this week SAP kept quietly building out the tools that let other companies build their own version of that coworker.
The clearest sign of that came from Joule Studio, SAP's platform for building custom AI agents, which now lets developers build those agents using Cursor and Claude Code — the same AI-powered coding tools programmers everywhere already use — instead of being locked into SAP's own editor. Once built, the agents run on a fully managed cloud runtime, meaning SAP itself handles the scaling and monitoring, so a company doesn't need its own team babysitting servers just to keep an AI agent online. Alongside that, SAP Domain Models — AI models trained specifically on SAP's own code, customer data and business processes rather than generic internet text — are moving out of early access and toward general availability this quarter, which matters because it's the difference between an assistant that knows AI in general and one that actually knows your company's specific setup. Joule also picked up the ability to span multiple S/4HANA systems inside a single connected "formation," so a company running separate systems for different regions no longer needs a different login and a different Joule for each one. Smaller but very real quality-of-life upgrades landed too: AI-assisted Easy Fill and plain-English error explanations went fully available, meaning instead of a cryptic error code when a form gets rejected, an employee now gets an actual sentence explaining what went wrong.
None of this is a flashy demo moment — SAP even has partners running a formal adoption campaign through the end of September just to get more customers actually using what's already shipped, which tells you this is still very much a rollout in progress rather than a finished product.
The bigger pieces to watch are Joule Studio 2.0 and the new SAP AI Agent Hub, a marketplace for managing and governing AI agents across both SAP and non-SAP systems, both targeted for later this quarter. The real test for Joule isn't whether SAP can keep announcing capable-sounding features — it's whether all these pieces add up to something an ordinary employee actually notices helping them get through their day.
Einstein AI
This Week in Einstein AI
Einstein is the AI brain built into Salesforce, the software millions of sales and customer-service reps use every day to track deals, log calls and answer support tickets. It used to just suggest what to type next; lately it's been graduating into Agentforce, a system of AI agents that can actually go do the work themselves. This was the week that graduation went fully public, at Salesforce's massive annual Dreamforce conference in San Francisco.
The headline from Marc Benioff's September 16th keynote — shared on stage with Anthropic's Dario Amodei and Nvidia's Jensen Huang — was AIforce, a new interface layer that lets Salesforce's data, workflows and permissions show up wherever people already work, instead of forcing everyone back to the same Salesforce screen. Benioff called it an "interface revolution," and the point is blunt: your customer data and your AI agent can now live inside Slack, or inside Claude, instead of only inside a Salesforce tab you have to remember to open. Backing that up, Salesforce's Claude partnership deepened with Claudeforce moving into open beta, and the company detailed new plumbing underneath it all — pieces with names like Koa, Agent Fabric and Guardian — built to give agents guardrails so a company can actually trust one to take real action instead of just suggesting one. Just days earlier, on September 11th, Salesforce gave its agents actual names and job titles: Casey handles customer service, Paige handles IT and HR requests, Carter helps shoppers, Marshall runs supply-chain busywork, Piper chases inbound leads, and Fin covers customer experience, with a seventh, Hunter the outbound sales agent, still in pilot ahead of a November release. Six of the seven are already generally available, and customers can rename them to match their own branding. Salesforce also leaned on scale to back up the hype, citing billions of these agent-driven work units already processed across its customer base.
It's a lot to take in, and that's sort of the point — Dreamforce is where Salesforce makes its biggest bets visible all at once, and this year's bet is that the future isn't one app you log into, it's an AI layer that quietly follows you into whatever app you're already using.
The thing worth watching next is whether AIforce actually ships broadly and works as advertised, or turns out to be more keynote promise than daily reality — and whether giving AI agents friendly first names changes how much people actually trust them with real customer conversations.
🔍 Tool Spotlight
Midjourney
Midjourney is the AI art tool that made "just type what you're imagining" into a real way to make pictures — describe a scene in plain words and it paints something back, often good enough to stop you mid-scroll. For years its one weak spot was editing: it was brilliant at generating something new, clumsy at fixing what you already had. That changed with its Edit Model, opened for public testing on August 27th, which folds three older, separate tools into one and lets you change an existing image just by describing what you want different, while also handling inpainting, outpainting, and blending up to four reference images at once. A September 16th update sharpened the experience further: the prompt bar now directly asks "What would you like to change?" when you attach a photo, draft batches show your source images in a clean grid with hover previews, a genuinely useful "Enhance" button replaced an upscale tool nobody liked, and Korean language support arrived for a huge new pool of users. Why it matters: this turns Midjourney from a one-shot image slot machine into something closer to a real creative tool, where you can keep refining a picture instead of starting from scratch every time you don't like a detail. Illustrators, marketers, indie game developers and anyone who iterates on visuals for a living stand to benefit most, and the surprising part is how much got simplified — three confusing features became one that just listens to what you say.
Source: Midjourney Alpha Changelog / PowerDrill AI, September 2026
AI Developments — Friday, September 18, 2026
ChatGPT
This Week in ChatGPT
ChatGPT is the chatbot everybody's heard of — the one you type questions into and it types answers back, except lately it also books things, browses the web and writes whole documents for you. Last week the big story was a brand-new brain, GPT-6 Astra. This week OpenAI turned its attention to something less flashy but just as consequential: how the company that runs the world's most popular AI assistant actually plans to make money from it, and how it keeps the plumbing underneath from getting confusing.
The headline move landed on September 16th, when OpenAI announced Sponsored Agents, a test where clicking certain ads inside ChatGPT doesn't just take you to a website — it opens a live, clearly-labeled conversation with an AI agent working for that advertiser, so you can ask it questions, push back, and only click through to buy something once you're actually convinced. It arrived alongside new tools that let businesses build ad campaigns just by describing what they want in plain English, plus fresh integrations with HubSpot and Shopify. The context makes it land harder: OpenAI's ad business has apparently already hit a $1 billion annualized run rate in roughly 200 days, running in more than 40 countries, so this isn't a side experiment, it's a company leaning hard into advertising as a real revenue engine. A day earlier, on September 15th, OpenAI said it would retire GPT-5.5 from ChatGPT, ChatGPT Work and Codex on October 14th, nudging everyone still using it toward newer models like GPT-5.6 Sol or the new flagship GPT-6 Astra — a reminder that in this industry even a model that launched back in April can suddenly look like old furniture. And on the more mundane end, ChatGPT Business admins quietly got a tool called Model Test on September 11th, which lets IT teams check exactly which models and settings a specific employee can actually access, without changing anything, just so troubleshooting "why can't Sarah use this feature" stops being a guessing game.
None of these are the kind of update that makes you gasp, but stack them together and a pattern shows up: OpenAI is simultaneously building out a genuine advertising business, cleaning house on old models nobody needs to be running anymore, and giving the corporate customers footing much of the bill better tools to manage who gets what. It's the unglamorous work of turning a wildly popular product into a durable company.
The thing worth watching next is simple: does Sponsored Agents stay a small, clearly-labeled experiment, or does it start creeping into more of the conversation? An AI assistant that's also quietly selling you things is a genuinely new kind of relationship, and how OpenAI draws that line will shape how much people trust the assistant with everything else.
Mistral AI
This Week in Mistral AI
Mistral is the French AI lab that's spent the last three years proving Europe can build serious, open AI models instead of just renting American ones. Last week the story was money — a record-breaking €3 billion funding round. This week the story was about where Mistral's technology actually shows up in your life, and the answer turned out to be a browser most people already have installed.
On September 16th, Mozilla announced that Firefox's new AI assistant, called Smart Window, is now powered by Mistral's models as part of a formal partnership between the two companies. Smart Window is the tool that helps you make sense of a messy search, remember something you clicked away from earlier, or pull useful context out of the tabs you already have open, and it's rolling out first to users in France and North America, with the UK and Germany expected to follow later this year. The privacy angle is really the whole point here: conversations aren't saved on Mozilla's servers by default, and Mistral agreed to zero data retention on its end, which matters a lot to Mozilla's brand as the browser that doesn't spy on you. Mistral's CEO Arthur Mensch called it a partnership between "two open source advocates," and it's a smart pairing — Firefox gets a genuinely private AI feature to compete with Chrome and Edge, and Mistral gets its models sitting inside a browser used by hundreds of millions of people who never signed up for a Mistral account at all.
Underneath that headline, Mistral's engineers also kept polishing the developer-facing tools that don't make front-page news but matter to anyone actually building with Mistral's models: the open-source Mistral Common library picked up better support for chat templates, cleaner handling of images and audio mixed into the same conversation, and fixes for running the tokenizer offline, the kind of housekeeping that makes life easier for the programmers stitching Mistral into their own products.
Compared to the funding fireworks of the week before, this was a quieter, more distribution-focused stretch for Mistral — less about raising money and more about making sure Mistral's AI is actually sitting in front of ordinary people, in a browser they already trust, rather than locked away in a developer's API console.
AI Developments — Thursday, September 17, 2026
Anthropic Claude Code
This Week in Anthropic Claude Code
Claude Code is Anthropic's AI that lives right in your terminal and actually does the work — reading your files, running commands, writing and fixing real code — instead of just chatting about how it would do it. This week's changes weren't the kind that make headlines; they were the kind that make a company's engineering lead trust the thing enough to let it near production.
The headline release, version 2.1.269 on September 11th, added a command called claude plugin eval, which lets developers actually test whether a plugin they built helps or hurts. It runs each test case three times with the plugin switched on and three times with it off, then scores the difference and spits out a report in both JSON and a browsable HTML page, so "trust me, it works" gets replaced with an actual number. The same release gave the VS Code extension an "agent map," a visual layout of every subagent Claude Code has spun up to divide a big job into pieces, plus a new setting that lets teams cap how many of those subagents can run at once — anywhere from just 1 up to 256 — so a company can dial up parallel work without accidentally melting their own infrastructure.
A handful of smaller fixes rounded things out. A day earlier, on September 10th, a maxEffortLevel setting arrived that lets teams cap how hard the model "thinks" on every single request across Bedrock, Vertex and Foundry, which is really a cost control dressed up as a feature. Mid-week brought new header information for companies routing Claude Code through their own servers, smoother handoffs when you're controlling a session remotely from the Claude app, and a clear warning the moment an MCP tool connection drops instead of a silent failure. Probably the most quietly beloved fix: a bug that made harmless, read-only commands like "git status" repeatedly nag you for permission during long sessions finally got squashed.
None of it is a single "look what it can do now" moment, and that's sort of the point — Anthropic spent the week tightening the bolts, the unglamorous kind of engineering that decides whether a team actually lets an AI agent touch its real codebase unsupervised.
Xai Grok
This Week in Xai Grok
Grok is Elon Musk's chatbot, wired directly into X and known for a looser personality and a tight relationship with Musk's other companies. The week's biggest story, though, was about a model that got promised and then quietly didn't show up on schedule.
Back on September 2nd, Musk announced Grok 4.7 would land in "about ten days," built with 2.1 trillion parameters — a 40% jump over Grok 4.6's 1.5 trillion — and trained partly on SpaceX's own rocket, satellite and manufacturing data to make it sharper at engineering problems. Then, right around the day it was supposed to ship, Musk said it needed a few more days. The reason he gave is a genuinely interesting glitch: during training, the model got penalized too harshly for writing long answers, so it learned to give up on hard problems early and skip double-checking its own work, a bit like a student erasing a correct answer because it took too long to reach it. Musk also admitted flatly that Grok 4.7 hadn't yet beaten its rivals, while teasing that Grok 5, still to come, is being aimed at something closer to true general intelligence.
While the flagship model slipped, xAI kept the rest of the shop humming. Grok Build, the company's coding-agent tool, shipped a steady stream of fixes — clearer diffs, smarter permission prompts, subagents that resume work properly instead of losing their place, a fixed terminal lag bug, and a cleaner unified dashboard. And from September 15th through 17th, xAI ran "Grok Bot Galaxy," a three-day livestreamed experiment in San Francisco where a small team tries to build an entire company from scratch leaning on Grok Bot, xAI's new autonomous "AI worker" product for businesses, complete with access and audit controls; Grok and Cursor Enterprise customers got two weeks of free access to try having an AI teammate run real tasks alongside them.
So Grok is stuck between two speeds right now: a marquee model that isn't quite ready to show its face, and a cluster of smaller agent tools getting real, visible polish. Whether Grok 4.7 actually lands looking as good as promised is the thing worth watching next, especially with Meta and OpenAI both expected to drop competing model updates before the month is out.