The weekly note
The week in agentic AI: July 20-25, 2026
A week in which the value chain moved at both ends at once: the model got cheaper while the lab climbed up to sell the deployed operation. Meanwhile, the AI Act sits one week away from its most serious date so far. I read it with an operator’s eye: what changes a decision, and what is noise.
What happened
Anthropic launched Claude Opus 5 at half price. On Friday the 24th, Anthropic shipped Claude Opus 5: capability close to its top model, Fable 5, at half the price ($5 per million input tokens, $25 per million output), with a three-level effort selector to trade cost against depth per task. On the company’s own benchmarks it beats the top model on 8 of the 13 published tests.
For a company running agents, what matters is the price. The “good enough for most tasks” tier of capability now costs half of what it did a month ago. If your architecture chooses a model per task, you capture that drop the same week, with an evaluation run and a configuration change. If your process is hardwired to one model, you don’t. The model offering for business has moved at this pace for months and shows no sign of slowing down.
OpenAI started selling the deployment, not just the model. On Wednesday the 22nd, OpenAI introduced Presence, a platform for deploying voice and chat agents in customer service, outbound sales, and internal support, with policies, approved actions, simulations, and evaluations built in. There is no self-service: it comes with a deployment led by the lab’s own forward-deployed engineers and selected systems integrators. As a reference point, OpenAI says its own English-language phone support runs on Presence and resolves 75% of calls without a human.
The move confirms where the market is forming. The lab has climbed a step up the value chain and now sells the deployed operation on top of the inference. It is the service as software thesis executed by the model maker itself. For the buyer, the useful question is the same as always: who keeps the operational architecture (the rules, exceptions, escalation, evaluations) and what happens the day you want to swap the model underneath. A vendor’s FDE is mandated to deploy their employer’s stack, and that shapes the answer from day one.
HubSpot opened its agent console. On Thursday the 23rd, Agent Hub and Agent Builder went into public beta for customers on the higher plans: one place to build and supervise sales and marketing agents that share the CRM’s customer context, with a no-code builder and credit-based billing when an agent executes actions. The problem it says it attacks is real, and I have seen it in live operations: a prospecting agent emailing an account the same week another agent is handling that account’s complaint, with neither aware of the other.
For a mid-market company, the reading is that sales agents now arrive embedded in software it already uses, with no implementation project. The trade-off is the cost model. Paying per executed action turns the agent into a variable cost you have to measure, and that cost belongs in what running agents actually costs, not just in the license line.
The AI Act reaches August 2 with its calendar split in two. One week out, the picture is this. Article 50 of the AI Act takes effect: disclosing that an AI is involved when someone talks to an agent, marking generated content, and labeling deepfakes. The Commission’s enforcement powers over general-purpose models also switch on, with fines of up to 15 million euros or 3% of worldwide turnover. And the omnibus package, approved by the Council on June 29 and pending only publication in the Official Journal, defers the high-risk obligations to December 2027 and August 2028. One nuance matters: the technical marking of content gets a four-month grace period (to December 2), but only for systems already on the market before August 2. Anything deployed after that date marks from day one.
What the AI Act means when you operate agents I cover separately. In short, transparency is serious and due now. High risk is deferred, not gone.
How to read it from operations
The pattern of the week is that the stack is commoditizing at both ends. At the bottom, frontier capability ages into utility within months, and every price cut confirms it. At the top, the lab and the SaaS climb up to sell deployment, control, and supervision. What sits in the middle, encoding your specific process with its rules and exceptions, remains custom work, and it is where the outcome gets decided.
What matters for the decision:
- Only companies that can swap the model capture the price drop. Opus 5 at half price is worth nothing if your prompts, controls, and evaluations are wired to another model. Model agnosticism is the cost lever, not a philosophical stance.
- If the model maker runs the deployment, negotiate the exit before the entry. With Presence, as with any vendor-led deployment, the contract has to say who owns the rules, the traces, and the evaluations when the agreement ends.
- AI Act transparency has run out of runway. If one of your agents talks to customers or generates content, the check is due this week. The marking grace period only covers what is already deployed.
What is noise:
- The benchmark race between labs. For your process, the benchmark that matters is your evaluation on your cases, with your data and your error thresholds.
- Counting embedded agents as digital headcount. An agent in beta that bills per credit is a variable cost per action with supervision still to be defined, not an employee already at work.
What to watch
August 2 is a double date: Article 50 transparency and enforcement powers over general-purpose models. One loose end remains, and it is the publication of the omnibus in the Official Journal, expected before that date, which is the act that makes the high-risk deferral legally final. If you run customer-facing agents, this week’s short list is concrete: a visible AI disclosure, marked content, and an inventory of which of your systems were on the market before the 2nd, because the marking grace period depends on it.