Death by verbosity: how to get concise AI responses
AI answers keep getting longer and we keep skimming them. Why models are verbose, what it costs, and a three-mode rule for concise AI responses.
I’ve been thinking a lot lately about the tempo and quantity of high-quality, productive conversations. One mode of conversation that has become very common has also become dysfunctional, and it has taken on a shape that I wouldn’t tolerate in other areas of my life.
I’ve surrounded myself with experts in all the crazy pursuits that I’m into in both my work and personal life. When I have a question in one of those areas, I’ll pick up my phone and call someone I’m reasonably sure knows the answer. The conversation (usually) follows a somewhat common pattern after we’ve progressed past the introductory niceties.
I always love injecting a little personal anecdote into each of these articles - and this will be no different. Today, we’re going to talk about pinball as my example conversation. Names have been changed to protect the innocent (or guilty!).
Me: Hey Jeff - I’m having a problem with my Star Trek: The Next Generation pinball machine and I’m wondering if you can help me.
Jeff: Sure thing.. what’s up?
Me: It restarts randomly - like, right in the middle of the game. It’s super frustrating.
Jeff: Does it happen more often when you hit both flippers at the same time?
Me: Yes! That’s virtually guaranteed to make it happen.
Jeff: Ah yeah, that’s a common problem with WPC machines. It’s usually Bridge Rectifier BR2 and/or Capacitor C5. It could also be cracked board traces to those components. Basically, the 5 volt rail sags on the board and the watchdog circuit restarts the game.
Me: Ah cool. I’ll try that and let you know if it doesn’t work.
That’s a nice short conversation that conveyed what the problem is nearly every time. I might have to look up additional information, or I could go on to ask Jeff more questions.
Lately though - every conversation has been more like this:
Me: Why is my machine rebooting?
Expert:
On Williams/Bally WPC pinball machines, unexpected reboots (or resets) during gameplay are almost always caused by a drop in the +5V power rail.
The CPU board features a watchdog circuit (specifically IC U10) that monitors this voltage. If the +5V logic line drops below ~4.7 to 4.8 volts, the watchdog immediately forces a system reset to prevent RAM/memory corruption. This typically manifests when heavy electrical demands load down the circuit, most famously when you press both flipper buttons simultaneously or during a high-action multiball.
While a low voltage triggers the reboot, the root mechanical and electrical causes can usually be traced to the following chain:
1. Connectors and Header Pins (The Most Common Culprit) Before replacing any major board components, experienced technicians look at resistive connections. Over time, Molex pins and IDC (Insulation Displacement Connector) plugs tarnish, loosen, or burn.
- J114 (Driver Board) & J210 (CPU Board): This specific path passes the +5V power from the driver board to the computer. Reseating or replacing these pins often eliminates the microscopic voltage drop that occurs under load.
- The Z-Connector: Many WPC games feature a structural “Z-connector” inline wire splice. This acts as a bottleneck where voltage tends to leak.
2. Failing Bridge Rectifiers and Filter Capacitors The power-driver board relies on a combination of bridge rectifiers and capacitors to turn alternating current into smooth direct current.
- Bridge Rectifier 2 (BR2) and Capacitor 5 (C5): On WPC-89 boards, BR2 converts the AC voltage and C5 filters it. If C5 dries up with age or BR2 begins failing, the “ripple” in the direct current increases drastically. When a flipper coil draws power, the weak capacitor cannot compensate, causing the voltage to dive.
3. Voltage Regulator Degradation The +5V rail is regulated by an LM323K (or similar) voltage regulator. If it runs hot over decades or the thermal paste backing it dries out, its ability to cleanly regulate high currents degrades, introducing dropouts under heavy loads.
4. Low Wall/AC Voltage Pinball machines are sensitive to line voltage. If your home or venue’s electrical outlet is outputting low AC voltage (e.g., dropping below 115V-117V), the transformer cannot step up enough juice to keep the secondary logic loops saturated.
5. Playfield Short Circuits or Bad Coil Diodes If a machine only reboots when one specific flipper or coil fires, the issue might be on the playfield rather than the boards. A broken or missing flipper coil diode allows a massive reverse-voltage spike to escape back up into the wiring harness, immediately knocking down the digital logic loop.
Popular Solutions
- The Diagnostic Approach: Re-pinning connectors, reflowing cracked solder joints on header pins, and replacing C5/BR2 or the regulator.
- The “Daughterboard” Fix: Many owners opt for a commercial WPC Reset Fix Board (such as the one from Marco Specialties). This bypasses the aging +5V rail entirely by taking the much stronger +12V rail and regulating it cleanly right at the CPU board.
At first glance, the second looks like a great answer - and it is - but without context, there are a lot of red herrings in there. Could it be low wall voltage? Yes, it could be - but almost all of us get good power from the grid. Spending time tracing down playfield faults can take forever and that’s time you’ll never get back.
The problem with the second response is that it’s too much information. As a one-off, it’s fine and complete. What happens when every single question you ask turns into a wall of information? It all quickly loses meaning because it’s too much to ingest and contextualize, much less act upon.
It’s pretty obvious who the second expert is - and anybody living in 2026 will recognize that I’d be foolish to deny the value AI lookups have added to my world. Can we have more human-shaped conversations that don’t make my eyes glaze over with too much detail? Humans just don’t talk like that.
I became more and more aware that I was skimming the answers to look for a nugget I was interested in. If I didn’t see it, I’d ask a follow-up and then I’d get another page (or more!) of text. This cycle would repeat to the point where I would have to scroll back many pages just to find the original answer - or more likely, I’d tell it to restate the answer because it’s too much work to find the original answer in the landfill of syllables. All of that text going by - and I absorbed almost none of it.
Why Is Every Answer a Research Paper?
None of this should be a surprise. That behavior is baked right in. When a model is trained, the lab shows pairs of responses and asks which one is better. Do that enough times and the votes train a scoring model. The language model then gets tuned to chase that score. Here’s the catch - more often than not, people voted for the longer response. One research team built a scoring model that ignored the words entirely and rewarded nothing but length, and it reproduced most of the improvement (Singhal and colleagues, 2024). We train our models to give longer responses and they dutifully comply to a fault. OpenAI already ran into a cousin of this problem in 2025, when it rolled back a ChatGPT update for telling users what they wanted to hear, and its own postmortem pointed at the thumbs-up data.
Imagine how exhausting a conversation with a friend would be where every utterance was answered with a lecture.
At a minimum, you’d stop asking questions - or you’d interrupt mid-sentence, or perhaps find a convenient excuse to visit the powder room.
.. so why do we tolerate and reward our newest advisors for the same behavior?
The trend is going the wrong direction as well. Compare the recent 5.x models to older ones and the newer ones come out no better, and in some cases wordier. A recent study (YapBench, January 2026) tested 76 models and found that GPT-3.5 Turbo, a 2023 model, scored best on terseness.
The cost of this verbosity is an interesting angle.. for API users, output is 5-6x more expensive than input. That sounds great until you realize the additional income isn’t there on the flat-rate plans. It’s in the model providers’ best interest to be terse if they’re losing money on every token. The verbosity survives in spite of the poor economics, which points straight back at the training.
What Is the Risk?
We’ve known for years that it’s dangerous to provide more information than is strictly necessary. One study looked at how doctors responded to 1.3 million alerts in their records system, and the odds of a doctor acting on an alert dropped about 30% for each additional alert in the same visit (Ancker and colleagues, 2017). About half of security teams admit they turn off their noisiest alerts entirely when they can’t keep up (Critical Start, 2021).
This is what AI’s walls of text are doing to you. You’re being drowned in content and your productivity suffers.
Can’t We Just Tell It to Be Brief?
It dawned on me that less might be more.. or, as the internet likes to say, can we get AI responses that tl;dr themselves?
(aside: tl;dr is internet slang for “too long; didn’t read” - after a wall of text, it’s common to just include a one-line summary).
Can we reshape responses to look more human?
There’s a problem with telling a model to just get to the point. Giskard’s Phare benchmark in April 2025 showed that instructing a model to “answer briefly” could actually increase hallucinations. Mostly that’s because rebutting a wrong premise takes more words. Even if you limit responses to a more reasonable average - you lower the ceiling on how accurate the answer can be.
What Does a Human-Shaped Conversation Look Like?
Let’s look back at my chat with Jeff. There’s an arc to the conversation, and it follows the “Storytelling 101” playbook almost perfectly. For those not in the know, Storytelling 101 goes like this:
- Act 1: Introduce the cat
- Act 2: Cat gets stuck in a tree
- Act 3: Save the cat
An introduction, a couple of short clarifying questions to focus the result, a little more back-and-forth - then a reasonably complete (but not over-wordy) answer.
I always have the freedom to ask more questions - and get more answers - or not. The point is that I’m not being flooded with every possible thing it could be.
Three Modes
Here’s the general rule: an interactive conversation with clarifying questions. The first real answer runs longer, and the follow-ups stay short. When I ask for the full response, give me the full response. When my question is built on a wrong assumption, take the words to say so.
For starters, I don’t think it’s a one-size-fits-all solution. I want it terse until I don’t. When I ask for a page of information, I want that full page. With that in mind, there are probably three different conversation modes.
First mode: Conversational. Ask a clarifying question or two, give a complete first answer, then keep the follow-ups short. Terseness matters most here.
Second mode: Content Creation. When I ask for an email, a document, or a page of analysis, give me the length that output needs. The terseness rule covers the conversation. What I asked it to produce gets whatever length that output calls for.
Third mode: Memory Mode. The model retains all appropriate detail in any notes it keeps for itself.
OpenAI’s own prompting guide already splits the first two modes, short for chat and long for research agents. The third one is the one I don’t see anyone talking about.
The Experiment
I’m running an ongoing experiment to see if reshaping the responses helps. I’ve included this in my “Instructions for Claude”, which gets fed to every chat:
IMPORTANT: Default to concise, high-information responses. Exception: the first response to a new topic should be thorough and complete; after that, drop back to terse as we iterate. Lead with the answer. Do not restate my request, add a preamble, recap, or explain basics unless asked. Follow up with clarifying questions rather than stating every possible outcome. Offer to expand rather than anticipating every possible question.
Override: if being brief means dropping a caveat that matters or letting a wrong assumption stand, take the words. Accuracy beats brevity.
Scope: this applies to conversation with me. When I ask you to produce something (an email, a document, code, a summary), use the length that output needs. It never applies to notes, memory, plans, or working files you keep for yourself; those should hold whatever information you need to do the job.
I’m one week into the experiment and so far it looks good - I’m still tweaking the rules in real time. I’ll report back after a couple of months with it (the same way I review my own Claude Code usage). I look forward to the day when Claude converses like Jeff.
If your team is trying to put AI to work without drowning in its output, that’s a problem we spend a lot of time on. Here’s how we approach AI work at Tarmac.