Next Free Webinar

AI agents are showing up across departments, in sales, finance, and operations, whether executives planned for it or not.

This 60-minute session is built for the people who have to make the call on where agents belong, what they're allowed to do, and how to know if they're working. On Tuesday, October 20, COO Josh Sullivan will cut through the vendor noise and focus on the questions that matter:

• What to fund: where agents belong and where they don't yet

• What to watch: what agents are allowed to do, and how to know if they're working

• What to hold back on, and how to say so when a vendor is pushing

Who this is for: Executives and department heads who will decide where agents get deployed, what they can touch, and what counts as working, before the team starts using them.

Tuesday, October 20, 2026 · 12:00 PM ET / 11:00 AM CT / 9:00 AM PT · Free live webinar with Josh Sullivan, COO at Kiingo AI.
Save your spot now →

This Week's AI Rundown

• Google announced Gemini 4 Argon on September 30, six days after the new head of Google DeepMind said Gemini 4 would arrive as soon as possible. Google reports 77.9% on DeepSWE v1.1, a benchmark of real-world software engineering tasks, against 74.2% for Claude Opus 5.5 and 74.1% for GPT-6 Astra, and a tie for first on a benchmark of fixing security flaws. The model can produce up to 1 million tokens of output in one reply, up from 64,000 (tokens are the units AI usage is billed in). Introductory pricing is $2 per million input tokens and $10 per million output, rising to $4 and $20 after the launch period. It is rolling out first to vetted cyber defenders through Google's Fairwind Program, then to paid API customers and AI Ultra subscribers, with no date given. Bloomberg reported ahead of launch that some inside Google doubt it matches Anthropic's and OpenAI's top models. (Google, 9to5Google, Yahoo Finance)

• Claude now works inside Google Docs, Sheets, and Slides. Anthropic released a Claude for Google Workspace add-on in public beta on October 6 for Pro, Max, Team, and Enterprise plans, installed once from the Google Workspace Marketplace. It opens as a sidebar next to the file, reads what is open, and edits in place: rewriting and restructuring in Docs, writing and fixing formulas and building models in Sheets, drafting and restyling slides in Slides. Google records each change under the user's name in version history, and the default "Ask before edits" mode shows planned changes for approval before anything is touched. Organizations that block Marketplace apps need an admin to allow it. A day later Anthropic released Claude Haiku 5.5, its small model for high-volume work like summaries, classification, and customer support, at $0.10 per million input tokens and $0.50 per million output for requests under 100,000 tokens, which Anthropic says works out to about 75% cheaper than Haiku 4.5 on average, and added monthly API credits to Max and Team subscriptions, from $100 a month for Max to as much as $500 pooled for Team, for running your own apps and agents. (Claude Help Center, 9to5Google, Anthropic, VentureBeat)

• Microsoft will switch on usage-based billing by default for new Microsoft 365 Copilot Business subscriptions bought through resellers starting December 1, a month later than first announced. The metered experiences include Copilot Cowork, the Work IQ APIs for custom apps, and GitHub Copilot Harness, which draw down Copilot Credits at one cent each. The default spending limit is 4,000 credits per user per month, which admins can adjust, so a 100-person company could spend up to $4,000 a month on top of its licenses before anyone changes a setting. The default-on change starts in supported markets and excludes Australia, France, Germany, India, Italy, and several others for now. (Microsoft, Yahoo Finance)

• OpenAI will retire custom GPTs on December 11 and move their builders to plugins, which bundle instructions, reference files, and connections to outside apps into one package that runs in both ChatGPT and its Codex coding agent. The migration tool carries over a GPT's instructions, knowledge files, and connected apps; custom actions, model choice, and sharing settings do not transfer, old conversations stay accessible but do not attach to the plugin, and OpenAI warns that "a migrated plugin may respond differently." Enterprise workspaces lose the ability to create new GPTs on October 26 and can apply to defer the shutdown to February 11, 2027. Separately, OpenAI will test visual ads beside ChatGPT image results in the US later this month, began rolling out a more visual ChatGPT with tappable buttons and editable graphs inside replies, and is reportedly raising at least $30 billion at a valuation of about $1.4 trillion. (OpenAI Help Center, Virtualization Review, TechCrunch, TechCrunch, Yahoo Finance)

• Twelve licensed CPAs averaged 37% on month-end close tasks that frontier AI models completed at or near 100%, in a study Mercor published October 1. The accountants, averaging five and a half years of experience, mostly took 30 to 180 minutes per task and scored anywhere from 0% to about 90%; Claude Opus 5 scored 100% on all 20 attempts and finished each in under 10 minutes, at $0.21 per rubric criterion met against $10.35 for the accountants. Mercor, which recruits domain experts to build AI evaluations, says the tasks measured detail-oriented instruction following on well-defined work, with no coworkers to ask, no accumulated job context, and none of the communication that fills the rest of an accountant's job. (Mercor) Deciding how much of this kind of work to hand to agents is exactly what the October 20 webinar covers.

• Anthropic committed $100 million to train 10,000 "Frontier Deployed Engineers" by the end of 2027, engineers who lead Claude rollouts inside companies from an idea to a system in production. The Claude Frontier Academy, announced October 2, runs like a medical residency: in-person training, simulated deployments, then a 12-week residency leading a real project inside the engineer's own employer. First cohorts come from Accenture, Bain, Capgemini, Deloitte, McKinsey, Morgan Stanley, Novo Nordisk, and Commonwealth Bank of Australia. A day earlier Barclays said more than 16,000 of its staff use a Claude-based knowledge assistant and Claude sorts about 120,000 client emails a day in its markets division, and on October 6 Anthropic expanded its startup program to a free year of Claude Team with up to five seats plus $1,000 in API credits for companies founded in the last five years or funded in the last two. (Anthropic, Anthropic, TechCrunch)

• California Governor Gavin Newsom signed 13 more AI bills on September 30. SB 947, the "No Robo Bosses Act," bars employers from relying solely on an automated system to discipline or fire a worker and requires human review, reviving a bill he vetoed in 2025. SB 951 requires employers to disclose when AI drives mass layoffs or relocations. SB 574 stops attorneys from handing core legal work entirely to AI, SB 1111 covers AI-generated digital replicas and impersonation, two bills set rules for clinical decision tools, and two amend the state's AI Transparency Act. (Office of Governor Newsom, CNBC)

• Regulators moved on both sides of the Atlantic. OpenAI said October 5 it will add an invisible watermark called textGrain to ChatGPT and Codex text for EU users over the coming weeks to meet the EU AI Act's transparency rules, which took effect August 2. The watermark nudges word choices in a pattern a detector can read, API developers anywhere can switch it on, and only approved researchers get the detector. OpenAI's own tests show detection falling from about 92% to 66% when 10% of the words are swapped for synonyms, and a missing watermark does not prove a human wrote the text. Two days later Google opened its SynthID checker to everyone, a site where anyone can upload an image, video, or audio clip to see whether it carries the watermark Google applies to AI-generated media, which OpenAI and Nvidia also support. In Washington, the FTC opened a probe into OpenAI, Anthropic, and other AI developers over unfair or deceptive practices and potential harms to consumers, and a senior official said the agency could issue formal demands for information. (TechCrunch, TechCrunch, ABC News)

• Utah let an AI start writing prescriptions. Under the state's AI regulatory sandbox, Nolla Health's app can prescribe topical acne treatments to Utah adults who complete a 10 to 15 minute questionnaire and upload a photo of their face, choosing only from a short list of physician-approved creams, with no oral drugs, starting at $4.99 a month. Physicians review every prescription before it is issued for at least the first 100 patients, then review after the fact for the next several hundred, then sample at least 10% of prescriptions each month. Nolla says it is the first company in the US cleared to let an AI issue an initial prescription; Utah's AI office says taking part in the sandbox is not an endorsement of the product. (SiliconANGLE, Bloomberg, Medical Economics)

• A software flaw found by an AI model was under attack within a day. Zach Hanley, a researcher at the penetration-testing firm Horizon3, used Anthropic's restricted Mythos model to find a critical login-bypass bug in Rejetto HFS, an open-source file server, by spotting that the server leaked random numbers an attacker could reverse to forge an administrator login. VulnCheck detected exploitation from an actor in China the next day, and the fix is version 3.2.1 or later. JPMorgan CEO Jamie Dimon told Bloomberg that cyber risk has gone up tenfold since Mythos, and on October 6 Anthropic expanded its Cyber Verification Program that gives vetted defenders, red teams, and critical-infrastructure testers access to Mythos 5.1 with fewer safety blocks; it says partners found at least 129,000 verified vulnerabilities between April and July. (The Register, Bloomberg, Anthropic)

What Studies Are Saying

• Bain's Technology Report 2026 finds companies that treat AI as a full business transformation are posting 10% to 25% EBITDA growth. Those leaders spend four dollars on people and process for every dollar on technology, while as much as 90% of enterprises remain focused on deploying tools for narrow use cases. (Bain & Company, Sept. 29, 2026)

• PwC's survey of 49,364 workers in 48 countries found 64% used AI at work in the past year, up ten points, and 22% now use it daily, up from 14%. Daily users are more confident in their job security, 68% against 57% for occasional users, and 21 points more confident about learning new skills. (PwC, Sept. 29, 2026)

• Gartner's survey of 1,303 respondents at companies with $50 million or more in revenue found the organizations that track AI returns in a disciplined way report positive returns on 81% of their initiatives. They manage AI as a portfolio and drop underperformers, and 85% of functional leaders plan to raise AI spending in 2026. (Gartner, Sept. 1, 2026)

AI in Practice: The Context Audit

A long conversation with an AI assistant carries everything that was ever said in it: the brief from message three, the direction you reversed in message twenty, the assumption it made that you never corrected. The replies drift a little at a time, and from inside the chat it is hard to tell whether the model got worse or the conversation got heavy. Before you rewrite the prompt, find out what the conversation is already carrying.

1. Paste this into any chat that has started to feel off.

"Before we continue, list every instruction and assumption you are currently working from in this conversation. Sort them into three groups. STILL TRUE: I said it and it still applies. POSSIBLY STALE: I said it earlier, but a later message may have changed it. INFERRED: you concluded it from context and I never actually said it. One line each, no explanations, and tell me which group is largest."

2. Read the INFERRED group first. That is where the drift lives: the audience it decided you were writing for, the tone it settled into, the constraint it built out of one offhand comment. Most of it will be reasonable. One or two lines will explain exactly why the last few replies missed.

3. When stale and inferred outnumber true, start over on purpose. Open a fresh chat and paste in only the STILL TRUE lines, corrected where needed, as the opening message. Then give it the task that went sideways and compare the two answers. The difference is the baggage, and you just left it behind. This works the same way in Claude, ChatGPT, and Gemini, with nothing to enable.

Run it whenever a working session passes the point where you can remember everything you have told it. A STILL TRUE list that keeps coming back the same way is also the first draft of standing instructions, worth keeping in a Claude or ChatGPT Project or a Gemini Gem.

Note from Andy (Growth Marketing Lead @ Kiingo AI)

Think about a new hire’s first week. Sharp, well trained, and mostly useless until IT gets them logins. They can’t answer a customer question without the CRM, can’t check a number without the accounting system, can’t tell what’s in flight without the project board. Nobody doubts the person. They just can’t see anything yet.

That’s the state most AI assistants sit in at work. MCPs, short for Model Context Protocol, are the logins. It’s a standard way to connect an assistant to a system you already run, so for a given task it can look at the actual invoice, the actual ticket, the actual draft, instead of whatever you remembered to paste in.

I’ve been in the middle of a data integration project recently, and this is where it went from interesting to a blessing. The job was getting information out of several systems that had never been asked to agree with each other. With connections in place, I could have the assistant read from each one directly, show me where the records lined up and where they didn’t, and draft the fix, all in one conversation. The exporting, the side-by-side spreadsheets, the “wait, which version is current” went away. What was left was the part that actually needs a person: deciding what the data should look like once it’s together.

The interesting shift is that “what should I ask it?” turns into “what should it be able to see?” That second question is answered one connection at a time, and it’s a lot more concrete than the first.

Two numbers in this issue describe the same gap. Bain found the companies seeing real earnings from AI spend four dollars on people and process for every dollar on technology, and Anthropic just put $100 million into training the engineers who make deployments stick. The model was the easy part to buy.

Becoming AI native is that people-and-process work: information your tools can reach safely, rules for what they may do, workflows rebuilt around the capability rather than decorated with it, and a team trained to brief and check the work. Kiingo AI helps companies get there, whether that starts with hands-on training, a deeper build, or something in between shaped to your situation.