Back to Blog

Is a Good Puzzle the Same Thing as a Good Question?

Edward Roske

Dawn and I have each done a bit over 300 escape rooms, which is a number I’ve mostly stopped saying at dinner because of the face people make, and this week I turned the home page of my own website into one, so we’re well past worrying about the face.

If you’ve never done a room: you pay somebody about $35, they lock you in a themed basement (basement is not always a metaphor) for 60 minutes, and you and 3 friends find out some things about each other. That’s the pitch. It sounds terrible written down. It’s one of the real pleasures of my life, and Dawn and I are both TERPECA nominators and voters, which means every autumn we sit down and argue about rooms we played on different continents and try to rank them (terpeca.com, if you want to see what the rest of the world thinks is good). We keep separate counts, which has caused exactly the argument you’d expect, and neither of us will round.

I should get ahead of the LinkedIn post where somebody says escape rooms are secretly corporate training. They aren’t. Go do one anyway.

Some rooms are hard because they’re hiding things from you. There’s a key taped under a drawer, and no amount of thinking produces the key, because thinking isn’t the mechanism, searching is. Those rooms reward whoever on your team has the fastest hands and the least patience, they’re perfectly fine, and I couldn’t tell you a single detail about any of them now. The best rooms I’ve been in do the opposite, which is the thing I can’t let go of: they put everything in plain sight and then wait. What’s missing is a question precise enough for the room to answer. You get a desk and a filing cabinet and a wall of framed photographs and no instruction of any kind, and the natural human move is to walk around touching all of it, and the room just lets you, patiently, for 8 or 9 minutes. Then somebody stops touching things and says “what’s different about these photographs,” and the whole thing comes apart in about 90 seconds. The information had been sitting there in the light the entire time, being asked nothing in particular. (Edward, you have now explained this exact phenomenon to at least 6 people this year, 2 of whom were seated next to you on a plane and had nowhere to be, and you have never once noticed them checking the time.)

I’ve been chewing on that for a couple of years and I finally worked out why it nags at me, which is that I watch the same 8 minutes happen in conference rooms, with an AI.

Somebody asks a model a question, gets back something technically correct and completely useless, and concludes the model is stupid. Sometimes it is! Usually the question just had no edges on it. “How many books have I written” is 15. “How many of those can somebody actually buy” is 13, because 2 of them were written for training classes and never listed anywhere (the 13 are here, the other 2 you’d have had to take my class for). Both are true, and only one of them helps if you’re standing in front of Amazon with a credit card. The model will hand you either one with precisely the same confidence, which is the part that’s new, and the part I haven’t worked out how to be relaxed about.

A human expert saves you from this without being asked, and I don’t think we ever gave them credit for it. You say “how’s the forecast looking” to your VP of FP&A and she does not answer that question. She answers the one you meant, which is usually “what changed since last week and do you need to care,” because she’s sat in 200 meetings with you and she knows. A model has sat in zero meetings with you. It takes what you said at face value, answers it completely and fluently, and hands it back sounding like the most confident person in the building. So the whole burden moves to you, permanently. For about 30 years the hard part of my job was getting the data out (I wrote 15 books largely about getting the data out, most of them with Tracy McMullen, which should tell you how hard it was), and the data shows up instantly now, so the only hard part left is knowing what to ask of it, which is a worse problem than it sounds because I’ve never once seen it taught, and I have yet to meet the executive who thinks he’s the weak link.

So what makes a question good? Boundaries, as far as I can tell (I’ve gone looking for a more interesting answer and there isn’t one): who’s asking, over what window, counting what, excluding what, and what answer would change your mind.

“Is it accurate” has no boundaries and every vendor alive can answer it. “What does it do when it doesn’t know” has boundaries, and a good number of them can’t answer it, which is itself an answer. “Can I see a demo” gets you something rehearsed. “Show me a run where it was wrong and what you changed afterward” gets you something real, or it gets you a subject change, and honestly I’d take the subject change.

The other half is the part rooms teach much better than any meeting ever did, which is that you have to be willing to look stupid for the first 8 minutes. Every good room has a stretch where the single most useful thing available to you is to say out loud, in front of people you’d quite like to impress, “I have no idea what I’m looking at.” Say it out loud and you get out. I have watched the other kind spend 40 minutes confidently rearranging objects, which looks a great deal like progress and produces none of it, and I’ve been on both kinds of team in both kinds of room, and the boardroom version costs considerably more than $35.

Which is why the site is a room this week. You’re locked in with a system called VERBATIM that answers precisely what you asked and not one syllable of what you meant (I wanted to call it HAL, and was talked out of it by an AI, and no, I don’t feel great about that). Ask it how many locks are on the wall and it says “a quantity,” because you didn’t specify whether to count the ones still in the crate. It’s never wrong, which sounds like a feature until you’ve spent 20 minutes with it. Three questions with edges on them open the door, and behind the door is the actual list of 6 questions I use on AI vendors, each one next to the lazy version that gets waved through. The whole thing takes about 90 seconds, every lock has a hint button, there’s an answer key at the bottom, and nobody is getting trapped on my home page.

Dawn hasn’t tried it yet. I’m aware she’ll beat whatever time I think is respectable, and I’ve made my peace with it in advance.

Anyway. Go play the room, tell me if the third lock is too easy (I suspect it is), and if you’re in San Juan on October 9 and 10 come to the Caribbean AI Summit and find me, because I will happily talk about this for longer than you want. If San Juan is a reach, which it is for most of you, just write me at Edward@Roske.AI.

Asking good questions, Edward