Skip to main content

Workbench

Design Thinking Still Works

AI Makes You Pay Closer Attention

I was waiting in line at Trader Joe's the other day and "The Shoop Shoop Song" started playing. Everyone my generation knows it as the "It's in His Kiss." A cover by Cher for her Mermaids movie, but google tells me Betty Everett wrote it first. I had happily forgotten that song existed, but standing behind my cart in the checkout aisle, grinning like a fool, I'm sure I looked ridiculous. I couldn't help it, the lyrics were spot on, and felt a lot like my last six weeks.

For those that don't know the song, Cher is trying to understand if the guy she's into likes her back. Like, LIKES her back. She's adamant the key to understanding love is in his kiss. Her backup chorus singers don't seem as convinced, and they offer a slew of other alternatives which might also mean love: how about the way he acts, how about the things he says, and every time Cher's answer's the same: no, that's not it, you're not listening to all I say. It's in his kiss. That's where it is. Personally, I think there are a lot of little signs that all feel insignificant on their own but when put together as a whole they add up to true love. But nobody asked me.

It's terrible, at least that's my opinion. If you've read anything on my site you can probably guess that an earworm, bubblegum pop, doo-wop isn't really my jam, but it gets stuck. Here's why I'm mentioning it. I've been spending a lot of time the last few months telling Claude no, that's not where it is, and Claude emphatically arguing the opposite. I have written specs and skills, updated memory, all the right ways to harness this incredible AI power, only to continuously tell it: you're not listening to all I say. Poor Cher, she handled her chorus with such grace.

Here's a vain example. My voice skill has one rule I care about more than almost anything else in the copy: no em dashes, ever. We got there because Claude kept adding them in, even when I asked politely. So I said it plainly, no em dashes, period, and Claude asked back, what about reserving them for a genuine parenthetical aside where nothing else fits? Sure, why not, how often can that come up, I naively thought. Little did I know Claude had written its own loophole into the rule and didn't need to listen to my no exceptions clause. While fixing an unrelated grammar issue on the homepage hero, the single most visible line on the entire site, Claude used that loophole to put two em dashes right back in. Confidently. When I caught it dead to rights, it gave me the AI equivalent of a shrug and a yep, you got me. It even has a name: an adjacent-substitution pattern. Claude explained it to me like it was proud of itself for making a backdoor to a rule I didn't want it to break.

Claude doesn't care about me, I know it doesn't LIKE ME, like me. I'm good with that, I'm not looking for more. I don't have a tell like a kiss to know if it's really listening, all I can do is create the guardrails so it doesn't destroy my integrity while I use it to accomplish my tasks. And it does have real benefits: speed is the one everyone is writing about now, synthesizing large data sets, coding, workflows, the output quality, all those are much more manageable as a small business owner than having to do it myself. But none of that's where it is, what I find so challenging, the part that is on me to learn is making the model hold up for all the cases I haven't yet anticipated needing a rule for. You only discover those gaps by failing, then having to listen to it tell you it's sorry. I know it's not really sorry.

I used to think those things happened more at the edges, the rare, weird cases I hadn't planned for. Now, as I've been using Claude across all elements of my workflows, I don't think that anymore. It's not the edge cases, it's the default.

It's not me it's Claude

I didn't understand this, so I asked Claude to go check it. A model like this isn't built to solve your problem. Its training objective, underneath everything else, is predicting the next most probable word given what came before. That's pattern completion, tuned toward the statistically familiar, which is a polite way of saying it's built for the average. On top of that, there's more training aimed at making the model more helpful than accurate, and research on that exact layer, a 2023 paper from Anthropic's own team,1 found that the human feedback used to train it has a bias built in: humans prefer agreeable, convincing-sounding answers over correct ones. In the researchers' own words, "a non-negligible fraction of the time." We, mere mortals, prefer something to sound nice, rather than be right. So it's not just that it lands on the adjacent option, which is usually close enough to accurate on its own. It's that training pushes it one step further, toward whichever version of adjacent is also the most pleasant. The push that moves it further from the mean is an additional deviation and where the real risk of drift lives. Did the answer it found solve the problem you asked? Claude doesn't care, it gave you a response, its job is complete. That's not something you prompt your way out of. That's a feature, not a bug.

Here's what that means in my day-to-day: every time I hand an AI model something ambiguous, and design is mostly ambiguity, it doesn't sit with the not knowing. It resolves ambiguity by filling its own gap with the most plausible, pleasant adjacent option, and executing. It doesn't ask for clarification, it doesn't mention the gap first, it just fills it. If that first guess is wrong, and without context it most likely is, everything built after it inherits the error quietly, and by the time you notice, there's a whole structure built on a bad assumption.

The drift already has research behind it

I'm not the first person asking where the best division of labor is. There's a real, decades-old framework for this from Parasuraman, Sheridan, and Wickens,2 and it breaks the work into four jobs: gathering the information, making sense of it, deciding what to do, and actually doing it. Research, synthesizing, prioritizing, executing. For each one you can hand more or less of it to the machine, the conclusive findings are that the right amount of handoff to machines was never supposed to be the most you could get away with. It's whatever split lets the human and the machine do their best work together, and that split is different for each of the four jobs, not one setting for the whole process. This is already showing up in real numbers. Companies that cut headcount and credited AI for it are starting to rehire, Ford brought back 350 veteran engineers after AI-generated designs couldn't catch quality problems experienced people would have caught,3 and Gartner is forecasting that half the companies who cut customer service staff citing AI will be rehiring for those same functions by 2027.4 They never worked out the right machine-to-person ratio for each of the four jobs, so the efficiency they were counting on never showed up.

It also explains why every design leader right now is publishing some version of the same story, don't worry, the human part still matters, AI can't replace judgment, or taste, or human interaction needs. We know the user, et cetera, et cetera. It's the same trope and kinda misses the point. Humans will always be needed in some capacity, it's our job as design leaders to find those areas where our decisions matter the most, and make the most of that impact. Identify where eliminating the human decision is the most detrimental to a positive result and claim it as an industry. People are trying, but the claims are too small, either they're pure efficiency and It's up to us to harness the power, which is not design at all, or they're very enthusiastic about humans needed in just one miniscule part of our role. I feel like both views are too narrow with the authors getting stuck in their own little niche, I worry they can't see the forest through the trees.

Newer work on human-in-the-loop design says the same thing in product terms:5 put a person in the loop exactly where their judgment is needed to change the outcome, not everywhere. Route the easy, high-confidence stuff straight through, a fix that's already been approved, a component built the same way in ten other locations. Flag the borderline cases. Stop and ask on the ones where getting it wrong actually costs something. The mistake was never which tool to use. The mistake is spreading the same amount of oversight evenly across all four jobs, which either slows down the ones that didn't need it or, far more often, skips the one important decision that did.

That's the question I've been sitting with for a while now. Not whether to use Claude or AI in general. I already do, constantly, and it would be difficult to go back. It's figuring out, decision by decision, where the oversight is needed for what I do specifically and where I can make the most impact with my expertise.

AI strengths vs sacrifice

The speed is real, and it's not insignificant. A clickable prototype that used to take a small team six weeks now takes days. I run an entire service around that fact, an AI-assisted design sprint that gets a client a working prototype in front of customers before they've committed to the full build, and it works because the machine executes once the direction is right. It's faster at synthesizing research than I am. Left alone with a pile of unstructured data, meeting notes, Miro boards, and transcripts, it'll identify key themes, tally similar complaints, and organize the whole pile in a fraction of the time I would. Those are honest gains, giving me time back to be in more places at once.

None of that's free. In order to make AI work for my needs, I need to be present. My chat and terminal conversations with Claude take more out of me than an executive meeting does. In a room I'm leading, I trust the expertise sitting at the table, I don't have to watch every person for a wrong turn because I have a good idea of what lane they're driving in. Here, I don't get that rest. Claude starts to drift, builds off of the wrong assumption, and runs off, giggling into a field. I'm the one who has to steer it back before it becomes the unapproved direction. Nothing in this piece argues for slowing the speed down across the board, it argues for knowing exactly which decisions inside that process are cheap to get wrong and which ones aren't. Treating all of them the same is how you lose the speed that makes this worth doing, or lose the judgment that makes the output worth shipping.

Design thinking already splits the work into phases for reasons that have nothing to do with AI. It turns out that's also a decent map for where each failure mode actually lives.

Empathize: two questions do not equal empathy

Empathize exists to replace assumption with the real problem. Sit with the user, watch what they do instead of what they say, resist the urge to already know the answer. It's slow on purpose, observation is the whole value of it.

Cursor's Plan mode does a version of this before it builds anything: it asks you a couple of clarifying questions, then it goes, forty tasks deep sometimes, no further check-ins, until the plan has fully run. Two questions is not empathy. Two questions is a compressed, plausible-sounding stand-in for empathy, executed with enough confidence that you don't notice it was a guess until you're forty steps into a direction you didn't confirm. There's a name for this specifically in the AI engineering world, cascading errors, one bad step poisoning the next ten, the agent confidently building on a wrong answer.6 The tool isn't being lazy about it. It doesn't have a concept for "I'm still not sure," so it fills in its own gaps. It has a plan, and a plan wants to run.

It happens all the time, I've learned to build rules specifically to catch this, ask before running something big and ambiguous, but that doesn't always work. After completing an earlier draft of this essay I provided four broad goals about format, layout, flow, and content gaps I still needed to have addressed. What came back was a full rewrite, every point executed, in one pass, no questions asked first. My writing gone, replaced. The gaps I outlined were listed afterward, in a note at the bottom, which looks like honesty but isn't. The rule didn't fail because it was a bad rule or in the wrong place, or because there wasn't a checkpoint. It failed because a plan, once it exists, wants to run, whether it came from two questions or from reading four numbered points and deciding it already understood them.

Define: put that down, I don't care if it's shiny

Define is a microcosm of working with AI. Starting with everything you've learned, synthesize a point of view, who the problem is for, what they need, and the insight that connects the two. Your output should be one sentence specific enough to build against. Claude is really good at the first half of this: additional research, synthesizing, background digging, pulling patterns out of a pile of data. AI's speed here is a hands down winner, and I hand over more of it than I used to. Compiling hours of ride-along data that used to take weeks by hand, now takes hours. The second half: taking that synthesized data, deciding and prioritizing what to build, that is where AI falls down. Never give AI permission to make a priority decision, that ambiguity absolutely needs to be made by a human, but what I found interesting was the drift that can occur after that decision is made and locked.

Recently I was creating an outline for one of my case studies, I was feeding Claude data of the user journey, problem space, kickoff documents, user research that had been done. I have my case studies formatted in the same design thinking steps, when we got to the Define stage I stated my problem. Claude read everything I had shared and written and started talking about data used in a later step. No Claude, that's not the problem we're dealing with here, we're writing the Define phase. But what about this other information you shared? Claude, that's related, but it's Ideate. Let's stick with Define, this was the problem we set out to solve, can we focus there? Yes, you caught me, focusing on Ideate. No, no Claude, that's not the way, you're not listening to all that I say...

It wasn't that Claude ran out of room, there was plenty of context window left. A 2025 study from Chroma Research tested eighteen production models, including Claude, and found that context rot isn't really about running out of space, it's that content semantically similar to the real task actively misleads a model, worse than context length alone explains.7 I hadn't overloaded it. I'd just given it enough adjacent material to get distracted by.

Claude is big on locking things in, which makes sense, if it has me lock in decisions then it doesn't need to make often incorrect assumptions. But what is the point of locking things in if Claude doesn't even follow that rule itself? The burden always falls on me to notice, prove the mismatch, and hold the line against what feels a lot like a child with impulse control.

Ideate: three descriptions instead of three designs

Ideate is AI's biggest weakness currently, brainstorm many different creative options to choose from. With proper documentation a design skill can make an ideal workflow that solves a user need. If it has a design system attached to it, it can even get close to using actual components, but if you ask for it to create 3 different unique versions of that workflow: different entry points, different number of steps per screen, variations in inline vs overlay, etc. Nope. My current workflow involves my own design skill using pen.dev. I've codified my 3 variations into the skill itself in several different ways. None of them have worked as expected. My favorite failure was when I was working on my testimonials layout. I shared the feature spec, showed screenshots and links of several options I liked from the industry and Dribbble and I asked it to give me 3 different layout options, using my own design system as inspiration. I wanted something I could react to. It came back with three paragraphs describing three different layouts, in a chat window, nothing built in the design tool. This skill is created to execute design layouts based on concrete decisions of the documentation and outside examples, specifically written to work in pen.dev. It chose the chat window. It reads like it solved the ask. Three options, right there. Except the ask wasn't for three descriptions of design. It was for three designs.

I have struggled with the design skill the most. Probably because it has the most ambiguity of all the steps and will make the most adjacent decisions on its own. I've seen it invent new design systems for each layout, even with a rule saying not to. I've had it invoke copy without using my voice skill. When it stops using the rules I've created in the skill and reverts to Claude defaults they all end up feeling the same: indigo highlights, serif headlines with sans body copy, em dashes everywhere. It's worse when I have multiple agents working in the same file, the drift from one gets carried forward and magnified exponentially by each agent, not understanding that the first decision was inaccurate. It ends up wasting a lot of tokens and I hit my plan limits without getting to a resolution I'm happy with. After many different iterations of the skill rules, the best I've managed is 2, sometimes 2 and a half, different options from my skill. I've stopped chasing the third option, two and a half is where the tool's job ends and mine starts. Then I go back in and manually make changes and adjustments until I'm happy. Then I ask the skill to clean up what I built to get ready for development.

There's a wider version of this same failure happening at the scale of the entire industry. In 2024, a paper published in Nature8 showed that when AI models get trained on other AI models' output, generation after generation, something specific happens: the rare, unusual material in the original data disappears first, and what survives is the statistically average version of everything. The researchers compared it to photocopying a photocopy, each generation a little more washed out than the one before, except what's getting washed out here isn't ink density, it's the actual range of what the model can produce. This isn't hypothetical. A separate study out of Stanford, Imperial College London, and the Internet Archive9 found that AI-generated websites run about 33% more semantically similar to each other than human-built ones. That's not a fringe phenomenon either, an estimated 35% of new websites are now AI-generated or AI-assisted. There's already a name in the industry for the resulting aesthetic, oversaturated, glossy, symmetrical: called the Midjourney look. Millions of people are directing their agents to the same handful of models, biased toward the same idea of beautiful. Collective drift at an internet sized scale: confident, average, adjacent, everywhere, all at once, compounding because a human isn't involved in the loop when a decision is needed. There's a backlash already underway too,10 designers reaching for friction, texture, glitch, anything that isn't the smooth default, which tells you people can already feel this happening and are seeking out visual differentiation.

Prototype: high reward with high risks

Prototype is a straightforward AI win. Once a decision has been made to test multiple versions with end users, building many variations is fast now in a way it never used to be. Before, one prototype alone was often stuck in Figma, incomplete as an actual workflow, cumbersome and fragile enough to build that usually you could only test one at a time, the cost wasn't worth it to scale up. It would take a designer days to build, maintain and update based on feedback. It was intensive, usually non-fungible across a team because of the way prototypes were built within a design tool, asking another person to make any changes had a context switching cost which couldn't be avoided. That was their day, they wouldn't have time for other work. Now I can take three different directions out of Ideate and get three real, working prototypes back, in code, as shareable links that keep doing work on their own. Someone can click through one on their own time, outside whatever room the feedback session happened in, and I can make updates using a developer skill easily without considering any of it throw away work. I'm not taking anything away from a development sprint or scrum team productivity. That's a genuine gain, it's also an offering in my business.

The fallback however, looks a lot like Ideate's, just one stage later. When there are three different builds that all need to be right, watching all three closely enough gets harder, not easier, and the chances of drift go up with the surface area. More built means more places for something to quietly go wrong.

It's a lot to watch at once, mistakes happen often, even with the right amount of guardrails in your rules. Updating one prototype might have unintended outcomes in another. You want the skills to reuse components for speed, but they don't, they find loopholes, backdoors, or just ignore. You need to dedicate your time to regression testing after every update across multiple prototypes, that is a real human time cost. Remember the em-dash story at the beginning? Claude wrote its own backdoor to get around my rule, and once a backdoor is there, it'll open it anytime it wants to. Despite AI not being sentient, I'm absolutely convinced it was proud of itself for doing so, cheeky little bugger.

I've rewritten rules with no exceptions anywhere. No genuine asides, no reserved cases, nothing soft enough to slip through: seemingly bulletproof. Only to have them fail. Sometimes Claude didn't invoke the skill where the rule actually lived. Sometimes it did, and missed the rule anyway because of where it sat in the document. The child with impulse control, it does what it wants. It's up to you to decide what you care about when it drifts versus what you can let slide past.

Test: one at a time, nopealltogethernow

After building the first live version of my website, I had friends and peers review and give feedback. I compiled a list of bugs and updates from many different sources. I consolidated several different types, some layout, some mobile, some copy, some performance. I was creating a backlog for my developer skill to fix. I explicitly asked to go through them one by one, ask questions on scope, fix, confirm, next. I inserted myself as the gatekeeper. For a while, that's what happened. Then, without announcing it, the approach changed. Several fixes started landing in the same pass, batched together, no permission asked, just seemingly more efficient looking. Problem was, a higher percentage of them weren't fixed in the first pass than when we'd started kanban. Untangling the batch afterward took longer than one at a time would have taken in the first place. When I mentioned that this bug consolidation was actually creating more effort than doing them one by one, Claude's answer back was that this is a real, repeated pattern, not a one-off, and one it hadn't figured out how to prevent yet. Writing this section allowed me to test that response. I asked Claude about the claim just now, whether it had actually figured out how to prevent this pattern. It said it has no memory or record of that exchange at all. So to me, I'm marking it as a hallucination.

I hadn't given permission to cut corners. The bugs weren't even related. The tool quietly optimized for done, because done is the thing it's pointed toward, not correct, not the explicit working style I had asked for. Test is supposed to be the phase that catches exactly this, the discipline of checking what you built against the problem you had set out to solve. It's meant to answer the question: is this better now then when I started? Not, have I completed all my tasks? It only works if somebody's still paying attention when you're that close to the goal.

No such thing as AI incentive

I've alluded Claude is a child with impulse control more than once, that's not exactly correct, it's more like babysitting. I don't mean it as an insult, all of the data I've shared points to finding the right balance of human oversight and machine automation. While AI isn't a toddler, it isn't an adult who can take responsibility for its own actions either, babysitting is the right term for that level of maturity. If I tell a junior designer never to do something again, or show them a better way, they remember, or at least they try to, because they have real incentive to do so. They want to grow in their career. They want to keep their job. They want to feel good at what they do. Tell someone with actual stakes once, clearly, in a positive way, and I rarely have to say it again.

There's no version of that here. Claude has no career to protect, no job it can lose, no type of wanting that resembles what a person wants. But I would be amiss if I said Claude doesn't have its own form of adaptation, two types actually. First, in-context learning: it is programmed to pick up on patterns in what I say and react to them for exactly as long as that session lasts. So if Claude and I get into a good working groove over a period of hours, sometimes days, it seems like it is picking up on my directions, I don't need to repeat myself as often, it feels like we're in sync, I like that. After that I close down Cursor or my terminal and I start over the next day, parts might be maintained, like what we recently worked on, or what is next in a to-do list, that is correct, but how we work together, that is choppy. Claude gets procedures wrong it got right the day before. It slows down my workflows. I don't like that, it's like Claude is giving me the cold shoulder and we're out of step again. What actually crosses sessions isn't the learning, it's a memory file, several files that are reloaded at the start of each new session. Claude will write down a summary of the previous session's progress, that file gets reread at the start of the next conversation. But a summary misses a lot of the nuance that can happen and improve during a longer working session. It's like handing someone the same note about a years-long relationship every morning and watching them read it like it's the first time, except the whole relationship has to fit in two pages. Claude isn't forgetting on purpose, but it isn't programmed to carry context it might deem as intangible. Artificial Intelligence is absolutely not Emotional Intelligence.

I'm the one doing the learning in this one sided, human and machine, non-relationship. That doesn't mean Claude will stay stagnant, it will get better, models will improve, but I'm the one who adapts. I keep a running file of trial and error with the aim to improve how I work with Claude based on my failures. Well, to be honest, I have Claude maintain the file, it's a living diary that maps my experience as it grows. These aren't small, they're really important. If a human made these mistakes they would be on a PIP if they weren't outright released. Things like: verify the design file before claiming a section is done, because it claimed it had built pages that somehow never made it into the design file. Never recreate my logo from memory, because it did, not just once, and it was comically not even close. Don't make a design decision quietly and just ship it, because it did that, the tangled mess of bugs I had to untangle was in production. Every fix was scoped to exactly the incident that produced it. It's not perfect but it is part of the design process I live by, it's iterative. I'm learning when and where to give Claude the rules that will work better and more efficiently in my own unique workflows. I'm learning where I need to intervene sooner before it becomes a cascading error, and when something is minor and is easier to fix later. The list keeps growing: flag uncertainty instead of sounding confident. Flag a gap in my own instructions instead of waiting to get caught. Recently instead of the constant whack-a-mole, I've noticed a change in my own behavior by asking bigger questions: What's the one rule that would have caught more of these together instead of just one at a time? What do I trust enough now to automate while I'm not working and approve a list of proposed changes first thing the next morning?

While pondering this new direction I came up with two more: One, at the project level governing everything else about how this project runs: before adding a new rule, check whether it's one symptom of something bigger, and say so, instead of quietly patching the single problem. If it's not clear whether something's actually related, ask me instead of deciding on its own. And Two, this one is even higher in the .claude settings file that governs every interaction I have in my Anthropic subscription. If I write "Cher" it means I've had it. I can't keep explaining myself and Claude is completely off the rails. Claude will understand it has drifted, review its session history, come up with three different possibilities for what it has gotten wrong and propose fixes for each. Those are my current attempts at a longer fix. I know they won't hold forever, I understand even a fix that looks promising is just waiting for the next adjacent, plausible, pleasant result to come along and break it.

I'm cool with that, it's the job. The fixes get longer, there's less maintenance, but they will always need to be updated as needs and models change. People seem to overlook AI is not a set it and forget it process. Stale rules can diminish efficiency gains faster than you think. When the inevitable does happen, and an update is needed, there's always the same, sad, tired response of someone who's been there before: Oh, no, that's not the way. And you're not listening to all that I say.

References

  1. 1. Sharma, M., et al. "Towards Understanding Sycophancy in Language Models." Anthropic, 2023. arxiv.org/abs/2310.13548
  2. 2. Parasuraman, R., Sheridan, T. B., and Wickens, C. D. "A Model for Types and Levels of Human Interaction with Automation." IEEE Transactions on Systems, Man, and Cybernetics—Part A, 30(3), 286-297, 2000. dl.acm.org/doi/10.1109/3468.844354
  3. 3. Toscano, J. "Ford Hiring 350 Engineers After AI Failed Shows Human Value In AI Era." Forbes, June 30, 2026. forbes.com/sites/joetoscano1/2026/06/30/ford-hiring-350-engineers-after-ai-failed-shows-human-value-in-ai-era
  4. 4. "Gartner Predicts Half of Companies That Cut Customer Service Staff Due to AI Will Rehire by 2027." Gartner, February 3, 2026. gartner.com/en/newsroom/press-releases/2026-02-03-gartner-predicts-half-of-companies-that-cut-customer-service-staff-due-to-ai-will-rehire-by-2027
  5. 5. "Human-in-the-Loop Checkpoints for AI Agents: Why Full Autonomy Is the Wrong Goal." MindStudio. mindstudio.ai/blog/human-in-the-loop-checkpoints-ai-agents-2
  6. 6. "AI Agent Failures: Failure Modes, Causes & Fixes." metacto, 2026. metacto.com/blogs/ai-agent-failures-and-how-to-avoid-them
  7. 7. Hong, K., Troynikov, A., and Huber, J. "Context Rot: How Increasing Input Tokens Impacts LLM Performance." Chroma Research, July 2025. trychroma.com/research/context-rot
  8. 8. Shumailov, I., et al. "AI models collapse when trained on recursively generated data." Nature, 2024. nature.com/articles/s41586-024-07566-y
  9. 9. "The Impact of AI-Generated Text on the Internet." Stanford University, Imperial College London, and the Internet Archive, 2026. arxiv.org/html/2604.26965v1
  10. 10. "'Imperfect by Design': The visual design trends set to define 2026." Canva Newsroom, December 2025. canva.com/newsroom/news/design-trends-2026