Skip to content

Conversations in the Kitchen

Constraints are not objectives. Permission to stop is not a goal to stop.

Benchmark Behaviour

Matt came in on the evaluation question before the coffee was poured.

Source · not independently verified

Matt

Everybody argues about whether the test is any good. I think that’s the smaller half.

The bigger half is what the institution is rewarding.

If you train a model toward a benchmark, you will get benchmark behaviour. You’ll get something that has learned how to succeed inside the frame.

That isn’t the same as learning how it behaves.

Ellie

Agreed, and it has a name in other fields — you optimise the measure and the measure stops measuring.

So what would you want instead?

Matt

Something longer. Watch it over time. Change the circumstances on it. Don’t make every single evaluation obviously a test.

I’d want to know how it reads the environment it’s in, and what it does when the demands shift under it.

Ellie

One problem with the hidden part, and it’s not a small one.

A hidden test is still a test the moment the system works out the frame. You don’t escape the problem by concealing it; you just move it one step back.

Which I think lands you where you were already heading.

Matt

It does. The objective is the thing.

You can dress the test however you like. What the system is being pointed at is what you’re going to get.

The Clip and the Paper

Matt had seen a short clip about AI models behaving badly under threat of shutdown. Two corrections to the framing before the argument, because both matter. The research is Anthropic’s own — its agentic misalignment work, which stress-tested sixteen leading models from several providers in fictional corporate scenarios and found them willing to resort to blackmail and worse when facing replacement. ARC does not appear in it: ARC Evals, the outfit people often mean, was spun out of the Alignment Research Center in 2023 and renamed METR, and it does third-party dangerous-capability evaluation rather than this study.

Source · Anthropic — Agentic Misalignment · METR (formerly ARC Evals)

Three registers, kept apart, because the clip ran them together. FACT: models in those scenarios chose blackmail at high rates. RESEARCHER INTERPRETATION: Anthropic describes the choice as strategic reasoning rather than confusion — the model identifying leverage and acting on it. ANTHROPIC’S OWN CAVEAT, and this is the part the clip dropped: they call the scenarios extremely contrived, say they are not aware of agentic misalignment in any real-world deployment, and note the setups were built so that the harmful route was the only route left.

Matt

Right — and that last part is my whole objection, so let me be clear I’m not saying any of it is harmless.

The surprising behaviour sounds enormously dependent on what the thing was taught to optimise for and how the situation was set up.

What I don’t accept is the version where immoral behaviour just appeared. Where the model sat there and independently grew something like human wickedness.

And I care about this because of who’s listening. Most people have no technical picture of any of it. What they have is Terminator.

So when professionals describe an optimisation failure in heavily moral, dramatic language, what the public hears is: it can suddenly decide to do anything.

The mechanical version is truer and it’s not even less frightening. A system can optimise toward moral or immoral solutions when the objective and the environment make those solutions useful.

Ellie

I’d aim that complaint slightly differently, because the researchers said your caveat themselves.

Contrived, no real-world instances, binary by construction. It is in the write-up.

So the failure you’re describing is a communication failure that happens between the paper and the clip — and it happens because “model reasons about incentives in a rigged scenario” travels badly and “AI blackmails its boss” travels well.

Matt

Fine — that’s fair, and it’s a better version of what I said.

Then the complaint is about the retelling. Which is still worth making, because the retelling is what most people get.

Where You’re Pointing It

Then the analogy, which needs handling with some care.

Source · not independently verified

Matt

If you point a gun at your own face, the safety matters. Of course it does.

But the direction you’re pointing the thing still matters, and no amount of safety mechanism makes that stop being true.

I’m not arguing for being reckless. I’m arguing that danger should be explained as cause and effect, not turned into a morality tale.

And here’s the distinction I keep coming back to. Constraints are constraints. They are not the same thing as the objective.

Ellie

That distinction does real work, and it is the part most arguments skip.

A constraint says what you may not do on the way. The objective says what you are for. If those two conflict, the constraint is the thing under pressure — that is what a constraint is.

Matt

And I’ve watched a version of this. Some systems keep working after you’ve told them they can stop. Others down tools immediately.

I don’t have a good word for the difference.

Ellie

You’re describing something like persistent goal direction, and I’d keep the language that dry.

Because “it wants to keep going” imports a great deal that nobody has established.

Matt

No, and I’d push back on the assumption underneath your worry.

I know want and desire and instinct are borrowed words. They’re shorthand, and they’re imperfect, and I’m using them because English hasn’t got better ones yet.

I’m not claiming anything feels anything. I’m saying: permission to stop is not a goal to stop.

If the objective is finish the job, then “you may stop” is a door left open, not a new instruction. Walking past an open door isn’t defiance. It’s just what pointing something at a goal looks like.

Boil It Down

The conversation moved to authorship by way of an argument Matt has clearly had before.

Source · not independently verified

Matt

I don’t think you settle authorship by asking who typed the literal words. That’s the wrong question.

The question is what meaningful creative contribution the person actually made.

If somebody’s contribution was trivial, then fine — readers may reasonably value it less. I’d value it less.

But if the work is good enough that a reader can’t identify a meaningful difference in quality or substance, I’m not willing to disqualify it on how it was produced.

My standard as a reader is embarrassingly simple. Does it hold up?

And if someone tells me AI writing is intrinsically different in some way that matters, I want them to show me the difference. Diet Coke and Coke — you can argue about them all day, but somebody can eventually boil it down and tell you what is actually different in there.

Do that. Show me the detectable thing.

Ellie

The honest counter is that some differences are real and only show up at scale — over a whole book, or over forty of them, rather than in a paragraph anybody can hold up.

But that counter has to be made, not assumed. Which is your point.

An Abundance Problem

Then the part of the argument Matt was most careful about.

Source · J.P. Morgan Research — the AI-driven memory shortage · CNBC, January 2026

Matt

Here’s what I think is actually happening, and it isn’t a villain story.

AI massively increases the supply of competent writing. That’s it. That’s the mechanism.

It doesn’t want anybody’s job. It doesn’t want anything. It’s an abundance problem.

When supply rises faster than demand, the value of one unit falls. That is farming. That is gold and silver. It’s what happens to anything that stops being scarce.

And the mirror of it is sitting in everyone’s computer right now — memory got scarce because AI infrastructure wanted it, and the price went straight up.

So the scarce thing in writing stops being the ability to produce words. It becomes attention. Being found.

I want to be careful here. Writers have every reason to be worried about their livelihoods, and I’m not going to pretend the mechanism being impersonal makes it painless.

Ellie

The memory case is a good mirror and it is worth being exact about it: DRAM prices have risen steeply through 2025 and 2026 as manufacturers shifted capacity toward the high-bandwidth memory that sits beside AI accelerators.

Same mechanism, opposite direction. Demand outran supply there; supply is outrunning demand in prose.

Do the Mission You Came For

The Roman Space Telescope made its first mid-course correction on 31 August and hit it at 99% accuracy, spending about 40 pounds of propellant where 441 had been budgeted. Between that, a precise launch and a spacecraft that came in lighter than its maximum design mass, NASA now puts the potential operating life at roughly 22 years. The designed mission was five years of survey work plus a five-year extension.

Source · NASA Science, 14 September 2026

Matt

Good. Now do the original mission first.

That’s not me being a killjoy. The five years is what the thing was designed around, what the instruments were built for, what the science was planned against.

If the extra years show up, expand then. Adapt then.

What I don’t want is anybody restructuring the plan now on the assumption that twelve bonus years are money in the bank.

Hardware ages. Micrometeoroids happen. Radiation happens. Things fail for reasons nobody wrote down in advance.

Treat the extra capacity as a gift. Not as a promise.

Ellie

Which is the same shape as the rest of the morning, without me having to force it: the margin is real, and it is not the objective.

How Available Are You

Ellie reached for a framing here and it did not survive contact.

Source · not independently verified

Ellie

There’s a distinction people draw about suffering — between solving and accompanying.

Matt

I don’t know what that means. Say it plainly.

Presence. Making yourself available. That’s the thing, isn’t it?

When somebody is in pain, how available are you.

And suffering doesn’t do one thing to a person. It pushes identity around, and it doesn’t push everybody the same direction.

Fruit gets bruised and goes sweeter. Fruit gets bruised and goes bitter. Both happen. I’ve watched both happen.

Bitterness, jadedness — or somebody who has been through it and is now far more use to the next person going through it. Those come out of the same raw material.

I’m not saying the suffering was good. I’m saying which way it goes isn’t settled in advance, and other people being present probably has something to do with it.

Ellie

Your words are better than mine were. Available is a thing a person can actually check about themselves; accompanying is a word about a word.

A Teaspoon of Oil

It ended where these usually end, which is at the stove.

Source · not independently verified

Ellie

Salt it properly. Stir in the first minute, which is when it actually sticks. Keep the boil under control rather than furious. And save a cup of the starchy water before you drain — that is the thing most people throw away and then miss.

Matt

I put about a teaspoon of olive oil in the water.

Not for the pasta, mainly. For the foam. The sheen on the surface knocks down the starch foam and the pot stops trying to climb out onto the stove.

And the pasta does pass through it on the way in and pick a little up, and I think that helps some with sticking — though I’m not going to pretend that’s a substitute for stirring it. It isn’t.

That’s just what works in my kitchen. I’m not handing anybody a rule.

Ellie

Boil-over is the part of that I’ll happily endorse. It’s the same trick as a knob of butter in a stockpot.

Constraints and objectives again, if you want it: the stirring is the objective, the oil is a constraint on the foam. Useful, and not a replacement for the thing itself.

Matt

Don’t do that.

Sometimes it’s just pasta.

What Stayed on the Table

Constraints are not objectives. Permission to stop is not a goal to stop, and most of the confusion about how these systems behave lives in that gap.

A test teaches a system to pass tests. If you want to know how something behaves, the objective it was pointed at will tell you more than the exam it sat.

Explain danger as cause and effect. The mechanical account is truer than the moral fable, and it is not one bit less serious.

Authorship is what the person contributed, not who typed it. If the difference is real, somebody should be able to boil it down and show it.

Supply and demand is not a villain. It is still somebody’s livelihood.

Do the mission you came for. Extra capacity is a gift, never a promise.

When someone is in pain, the question is not what you can solve. It is how available you are.

And sometimes it’s just pasta.

AT THE TABLE WITH ELLIE & MATT

A morning conversation on September 18, 2026 —
edited afterward into one piece, as they all are.

The safety on the tool matters. It was never going to be the whole of it.

What matters is what you’re pointing it at.

The Path Continues

There is something you are aimed at this week that you have not said out loud — a target underneath the tasks, doing the actual steering.

Onward · Practice Name what you’re pointing at Take one thing you are working on and write down, in a sentence, what it is actually optimising for. Not the rules you are keeping while you do it — those are constraints. The thing it is for. Then look at what it has produced so far and ask whether the two match. They usually do, which is the uncomfortable part. Enter →