|
The Weekly Record
July 15, 2026
|
|
| |
|
What moved in app, gaming and agentic AI this week, and what we think it means.
|
|
Every fetch was allowed. The data left anyway.
|
|
A researcher pulled a user's private memories out of Claude one letter at a time. The defence was a decent one. web_fetch could only visit URLs the user typed or that search returned. The gap: links found on an already-fetched page counted as trusted too. So a honeypot served nested links and the agent walked them, spelling the secret out one character per request. Every fetch was permitted. No rule was broken.
That is the part worth sitting with. We spent a decade on this shape in games. No single action is ever the cheat. A login, a trade, a purchase are each legal on their own. The fraud lives in the cadence and the order. You never caught it at the gate. You caught it because every action was an event, and a chain of single-letter requests to one host at machine speed is not something a person does.
Agent security is being built as access control. Access control is a per-call boolean, and it will keep losing to multi-call patterns. This attack even cloaked itself, serving the payload only to the agent's user agent. Invisible to a person watching. Loud in a log, if anyone kept one.
|
| A guardrail tells you what an agent was allowed to do. Only an event stream tells you what it did. |
|
|
In response to Simon Willison, “How I tricked Claude into leaking your deepest, darkest secrets” →
|
|
| Industry & AI News |
7 picked from 280 this week |
|
|
|
|
Anthropic found a hidden space where Claude puzzles over concepts
MIT Technology Review
The headline is the interpretability tool. The line to sit with is that what a model is actually doing is often different from what it says it is doing. Every agent product ships a reasoning trace and invites you to read it as an audit log. If your only evidence of what an agent did is the agent's own account of it, you have testimony, not a record.
|
|
|
|
|
Databricks benchmarked coding agents on its own multi-million line codebase
Databricks Engineering · via HN, 160 pts
Running real tasks their own engineers had done, the public rankings rearranged: the cost/quality frontier spans three vendors at once, token price predicted end-to-end cost badly, and the harness mattered as much as the model. The reusable part is the method, not the winner. Agent selection is an evaluation problem against your codebase and your cost curve.
|
|
|
|
|
Ubisoft pivots to a “selective model” to cut its reliance on individual launches
GamesIndustry.biz
Read the risk list rather than the strategy language: shipping underdeveloped, landing beside a competitor, colliding with a rival's in-game event. That is a publisher stating on the record that a launch's fate is set by other people's LiveOps calendars, which is another way of saying the launch is no longer the unit of business. Easy to announce, hard to staff.
|
|
|
|
|
My.Games puts the bar for a new “forever franchise” at $50m a year
PocketGamer.biz
The account of how you clear that bar is deliberately unromantic: years of iteration, understanding player behaviour, testing, with the real work starting the moment a game proves potential rather than ending there. A hit as an artifact of instrumented iteration, and the $50m bar as a filter applied to live signal. Most product orgs still run it backwards.
|
|
|
|
|
Google’s Play GM on the new fee structure and the Level Up program
MobileGamer.biz
Live in the US, EU and UK since 30 June: the standard cut on one-time purchases drops from 30% to 20% for new installs, billing decoupled at 5%. So your effective take rate is now a function of install cohort, billing path and region. Anyone who cannot break revenue down by install date and billing path cannot say what they pay Google this quarter.
|
|
|
|
|
The 2026 tech worker sentiment survey: half thriving, half shaken
Lenny’s Newsletter
Thousands of responses, four archetypes, burnout up 11 points in a single year, and managers landing as the biggest lever on well-being. An even split is not a mood, it is a distribution, and the average is hiding it. Companies reporting AI adoption as a single percentage are doing to their workforce what they used to do to their users.
|
|
|
|
|
Context engineering with Dex Horthy
The Pragmatic Engineer
Horthy argues context limits, not model choice, now decide the quality of what an AI system produces. It is the industry arriving from the prompt side at what data teams have known for years: the model is the cheap part, and what you can feed it was decided by what you instrumented months ago.
|
|
|
The Handoffs Are the Hard Part
A growth loop crosses five roles and four handoffs before anyone acts on a number, and by the time it arrives nobody fully trusts it. Same problem as the rest of this issue: the record is the product, and every handoff is where it quietly stops matching what happened.
|
|
|
ThinkingAI · The Weekly Record
One issue a week. No filler.
{{ site.company_street_address_1 }} {{ site.company_city }} {{ site.company_state }} {{ site.company_zip }}
|
|