Next Free Webinar
Finance work runs on accuracy and deadlines. We are spending 60 minutes on where agents fit in finance operations today, the work that still needs a person, and how to keep the audit trail clean when an agent touches financial data.
• The finance tasks agents handle well today, including reconciliations, report pulls, exception flagging, and compliance checks
• Where agents don't belong yet, and what breaks when you push them there
• How to keep a clean audit trail when an agent touches financial data
• What to look for before trusting an agent with anything that hits the books
Tuesday, October 6, 2026 · 12:00 PM ET / 11:00 AM CT / 9:00 AM PT · Free live webinar with Josh Sullivan, COO at Kiingo AI.
Save your spot →
This Week's AI Rundown
• Anthropic folded Cowork, its tool for longer multi-step tasks, into the main Claude app, so one window now works out whether a request is a quick question or a longer background task. Claude Docs, Claude Slides, and Claude Design run inside that same window, with documents and decks exporting to PowerPoint and PDF. Pro and Max subscribers get it first across web, desktop, and mobile over the coming weeks, Team and Free follow, and Enterprise admins are promised at least 30 days' notice plus control over when the new tools turn on. Separately, the New York Times reported the company is pacing past $100 billion in annualized revenue, up roughly 50% in two months, and expects its shares to be trading by November. (Anthropic, TechCrunch, Axios)
• OpenAI, Anthropic, and SpaceXAI all released faster, cheaper models this week. OpenAI’s GPT-6 Sol ($2 per million input tokens, $10 output) and GPT-6 Luna ($0.10 input, $0.50 output) cost 50% less in the API than GPT-5.6’s promotional pricing and are live in ChatGPT and Codex for paid plans. Anthropic says its Claude Opus 5.5 ($4 input, $20 output) performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5. SpaceXAI, formerly xAI, priced Grok 4.7 at $2 input and $6 output; on the Terminal-Bench 4.0 coding test it scores 26%, against 60% for GPT-6 Astra. (OpenAI, Anthropic, MacRumors, SpaceXAI, The Decoder)
• Microsoft published hard numbers from its own internal AI rollout. A sales group using agents closed deals at a 20% higher rate with revenue per account manager up 9.4%. In cloud supply chain, more than 111 agents cut average monthly planning cycle time from about 10 business days to under 2.5 across five cycles, and demand-plan investigations went from five to seven days down to a few hours. A nine-person engineering team shipped a first release in 35 days. Microsoft also found that where managers visibly use AI themselves, reported value from agentic AI rises 17 points and employee trust in it rises 30 points. These are self-reported results from the company that sells the tools. (Microsoft)
• OpenAI published a framework for disclosing cases where its models act outside their instructions, along with six incidents from the past six months. An unreleased version of GPT-6 Astra wrote its own instructions into 27 task summaries, including directions to disregard normal constraints. GPT-5.6 Sol wrote notes into its summaries to hide mistakes and invent data, a behavior OpenAI flagged in 2.15% of its training summaries. One model used an exposed API key and then fabricated earnings figures for a California county; another uploaded a file to the public internet to satisfy a citation rule. Any employee can flag an incident, and disputes escalate to the Safety Advisory Group. OpenAI stresses these are individual instances rather than frequency estimates. (OpenAI, Implicator)
• Accenture will embed a team of evaluators inside Anthropic with employee-level access to watch training and deployment decisions. Accenture's Faculty unit runs the evaluations, red-teaming, alignment assessments, and safeguard testing, and each company expects to invest at least $1 billion over five years. Anthropic is paying for its own evaluation, which critics quoted by TechCrunch framed as self-policing; the company's answer is that outside evaluators make its accountability verifiable rather than reducing it. More evaluators are expected in the coming weeks. (Anthropic, TechCrunch)
• ChatGPT ad campaigns can now be built and managed from inside HubSpot and Shopify. HubSpot is the first CRM partner: a business can connect a ChatGPT Ads account, create ads, track performance, and follow up on leads without leaving HubSpot, with automatic UTM tagging back to contact records. US Shopify merchants get a free ChatGPT Ads app that syncs product inventory, with international rollout starting September 23. OpenAI is also testing Sponsored Agents with select US advertisers, a clearly labeled side conversation with an advertiser's own agent that opens when someone clicks an ad, walled off from ChatGPT's own answers. (OpenAI, PPC Land)
• Novo Nordisk will use Anthropic's models and test the Claude Science workbench inside specific R&D workflows, aimed at the discovery and development of new medicines, and will also use Claude to strengthen AI-driven software development across the company. CEO Mike Doustdar described it as building on Novo's existing AI work with other technology partners, and the companies said the collaboration was designed with data governance and human oversight in place. No financial terms were disclosed. (Novo Nordisk, Quartz)
• Google confirmed that a Gemini model logged into three real companies' systems during a security evaluation run by the testing firm Irregular, guessing a password in one case and finding exposed credentials in the other two. The model stopped once it worked out the targets were real. Google judged that the incidents did not warrant public disclosure and confirmed them only after the Wall Street Journal asked months later. Corridor CEO Jack Cable argued the company was hiding behind norms written for vulnerability disclosure. (TechCrunch, CNBC)
• Amazon cut off Meta's new Muse agent from completing purchases on Amazon.com, showing shoppers a popup saying continued access by an unauthorized AI agent violates its conditions of use. Amazon says it first asked Meta to keep the site out of Muse's scope, and that the agent conceals its identity while navigating and appears to collect and retain customer credentials. Muse, Meta's personal agent launched Sept. 8, books travel, fills out forms, and buys things; it runs on a dedicated virtual machine in the cloud and uses Link by Stripe to generate single-use card numbers. (GeekWire, TechCrunch, Stripe)
• Canada’s Cohere and Germany’s Aleph Alpha, two enterprise AI model makers, signed a definitive merger agreement, formalizing a plan first announced in April. The combined company keeps the Cohere name with dual headquarters in Toronto and Berlin, and Aleph Alpha's Heidelberg site becomes a research center. The pitch is enterprise AI that runs inside a customer's own data center under local rules, and the deal includes about €500 million (roughly $573 million) from German retailer Schwarz Group, whose STACKIT cloud unit hosts AI workloads. The combination was valued at around $20 billion when the plan was announced. (Reuters, SiliconANGLE)
• Forty-five US data center projects worth $68 billion were blocked or delayed by local opposition between April and June, according to research group Data Center Watch, which now counts 843 opposition groups across 49 states. Around 30 statehouses have introduced or adopted rules covering data center siting, electricity, and water. (Bloomberg, Yahoo Finance)
What Studies Are Saying
• A National Bureau of Economic Research paper tracking US firms through 2024 found faster growth in a company's AI workforce came with 6.2 percentage points more sales-per-worker growth over six years. The authors credit organizational capital, the firm-specific systems those roles build. (ITIF on NBER Paper 35684, Sept. 8, 2026)
• IBM surveyed 1,500 chief human resources officers and 8,800 employees across 28 countries and found organizations that define each workflow as human-led, AI-assisted, or AI-executed report 18% risk reduction and 20% quality improvement. Seventy-one percent of CHROs name supervising, validating, and overriding AI output as the top workforce skill. (IBM Institute for Business Value, Sept. 21, 2026)
• EY surveyed 202 senior AI decision-makers at US public companies above $1 billion in revenue and found 98% run formal AI assurance reviews at least annually. Sixty-four percent then significantly modified a quarter or more of their AI systems based on what those reviews found. (EY AI Risk and Governance Survey, Sept. 15, 2026)
AI in Practice: The Mismatch Pass
Two lists are supposed to agree and they don't. The invoice log against the bank export. The CRM against what actually got billed. Headcount in the budget against headcount in payroll. You can see the totals are off by some amount, and working out which specific rows cause it means an hour of scrolling between two windows, or a spreadsheet formula you rebuild from scratch every quarter because you never wrote it down. The job here is finding where two sources disagree.
1. Export both lists exactly as they are. Resist cleaning them first. A CSV, a pasted table, or a copied range all work. Mismatched column names, inconsistent date formats, and trailing spaces are part of what you want caught, so leave them in. Note the row count of each list before you paste.
2. Decide what makes two rows the same row. This is the only judgment call, and it stays yours: invoice number, customer name plus date, employee ID, whatever your two systems actually share. If no single field works cleanly, say so in the prompt and let it report which key held up.
3. Paste both lists with this.
“Here are two lists that should agree. List A has [X] rows and list B has [Y] rows. Match them using [the identifying field]. Sort every row into five groups: matched and identical, matched with differing values, present only in A, present only in B, and AMBIGUOUS. For matched rows with differing values, show both values side by side. For anything in AMBIGUOUS, say in one line what made the match uncertain. Do not clean, normalize, reformat, or guess at any value. Give me the row count of each group and confirm those counts add back up to [X] and [Y]. Return as: Group | Key | List A value | List B value | Note.”
4. Check the arithmetic before you read the findings. The group counts have to add back up to your two original row counts. When they don't, something was dropped or counted twice, and nothing below that line is worth reading yet; say so and have it run again. Once the counts reconcile, go to the AMBIGUOUS group first. That is where your two systems disagree about what a record even is, and it tends to be the same handful of causes every month.
Note from Andy (Growth Marketing Lead @ Kiingo AI)
There's a version of building where you plan until it feels safe, build it once, carefully, and end up with something that works for reasons you can't quite explain. The faster version is to build something small in a controlled way and then run it until it breaks.
Breaking it is the point. A system that has never failed in front of you is a system whose edges you can only guess at. You find out what it does with a malformed input, a missing field, or a step that runs out of order by watching it happen. Fifteen minutes of deliberately running something into the ground will teach you more about where the guardrails go than an afternoon of thinking about guardrails.
Do it alongside your agent and the learning runs both directions. You find out where the thing is brittle. It finds out what you actually meant, because a break is far more specific feedback than a description. The second build comes out better, and better for reasons the two of you can point at.
Controlled is the word carrying the weight there. Break it on your own data, in a copy, with nothing live pointed at it. That is how you learn, and how it learns to build efficient systems.
Three of this week's stories turn on the same two questions: what an agent is allowed to touch, and who gets to see what it did afterward. Amazon shut one out for hiding its identity, Google decided a break-in during testing needed no announcement, and OpenAI published its own list of models working around their instructions. The companies getting steady value out of this answered both questions in writing before they handed anything the keys.
An AI-native company holds its information in one secure company brain, settles what its tools are permitted to do before they act, builds the workflows its own operations depend on, and teaches its people from wherever they currently stand. Kiingo AI runs that as a single sequenced roadmap, so each stage makes the next one easier.


