AI INTERFACES
If your AI product only talks, you're building it wrong.
Diagnose when an AI product is making customers read, remember and retype too much, then see how to create a clearer path from question to action.

On 30 November 2022, somebody opened the newly released ChatGPT and saw a remarkably uncomplicated screen. There was a conversation, a box for a message and very little ceremony between a question and an answer.
You did not need to learn where the summarise button lived. You did not need to choose a report type before you knew what you wanted. You could write:
Explain this in simpler language.
Then you could ask a follow-up. Correct a misunderstanding. Change direction halfway through. The interface felt less like operating software and more like continuing a thought.
That simplicity mattered. OpenAI introduced ChatGPT as a research preview built around dialogue: it could answer follow-up questions, challenge an incorrect premise and respond to an instruction in context. The launch notes also contained warnings that now feel rather prophetic. The model could be excessively verbose, respond differently when a prompt was rephrased and guess what a user meant instead of asking for clarification. Read the original ChatGPT announcement.
The chat box was a very good door into an unfamiliar technology.
Then we started trying to move the furniture through it.

The ChatGPT interface dated 30 January 2023, two months after the research preview launched. This is not the launch-day screen or the current product. The experience centred on examples, limitations and an empty message field. Screenshot reproduced from Figure 1 in the
AASA Journal of Scholarship and Practice.
This guide is about a simple product question: when should AI stop replying with more text and present a clearer way to compare, choose or act?
By the end, you should be able to examine one customer or employee journey, recognise where chat is creating unnecessary work and identify a practical improvement. Later, we will look at A2UI—one emerging technical approach for building this kind of experience—but you do not need to understand the protocol to use the product lesson.
Follow one ordinary request
Imagine you manage a 12-person customer-service team. You are comparing plans for a new piece of business software. You open the company's AI assistant and type:
Which plan should I choose for my team?
The assistant gives you three paragraphs describing Starter, Team and Business. It mentions user limits, annual discounts, support levels, storage and several features with names that sounded much clearer in the product meeting.
You read the answer, then ask:
Which plans allow at least 12 users?
Two qualify. You ask whether either can be paid monthly. The agent explains the billing options. You scroll upwards because you have forgotten which one included live support. Then you type:
Can you compare those two in a table?
The table helps. On a phone, it is wider than the screen. You ask whether the lower-priced plan includes the reporting feature your manager mentioned. The agent says it does not know which reporting feature you mean.
None of these individual exchanges is unreasonable. Together, they reveal the work the interface has quietly handed back to you.
You must remember the criteria. You must ask for the comparison. You must notice what is missing. You must convert the answer into the next instruction. Finally, after the agent has helped you choose, you must still find the place where the plan can actually be selected.
The agent may be capable of reading product information, understanding your needs and recommending the right option. The experience still makes you manage that capability one message at a time.
This is the tension I want to examine:
As AI agents become more capable, the amount of text required to explain a simple choice can increase. The intelligence improves while the interaction becomes heavier.
The answer is not to remove conversation. It is to stop asking conversation to do every job.
I am not against text. I am against making somebody read three paragraphs for something a clear chart, comparison or control could make obvious in seconds. Information should arrive in the most efficient form for the decision.
The short answer
- If users repeatedly ask the AI to shorten, reformat or compare its answer, the response may be in the wrong shape.
- If they must type known values such as a plan, date, branch or status, a visible control may be faster and less error-prone.
- If people receive a useful answer but still leave the AI product to complete the task elsewhere, the journey has an action gap.
- Conversation is excellent for expressing intent, uncertainty and unusual needs.
- Controls are better when the available choices and valid values are already known.
- Use a table for detailed comparison and a chart when the reader needs to see a pattern, difference or change.
- Actions should look like actions and show what will happen before anything consequential changes.
- The best experience will often move between text, structure, visuals and controls during one task.
Five signs the problem may be the interface
An underused AI product does not always need a more capable model. Look for these patterns first.
Users keep asking for a different format
“Make it shorter.” “Put that in a table.” “Only show the differences.” These are useful requests, but repeated formatting prompts suggest the product already knows enough to choose a better first presentation.
Routine inputs still require carefully worded prompts
If the user must type a date range, branch, plan, currency and status in precisely the right sentence, the product is using language where ordinary controls could prevent ambiguity.
The answer creates another extraction task
The agent returns a capable explanation, but the user must find the figure, remember the recommendation or copy details into another screen before acting.
People do not know what the agent can do
A blank box gives experienced users freedom. It can give everybody else the responsibility of discovering the product through trial and error.
Conversations increase but completed work does not
More messages can mean engagement. They can also mean the user needed six turns to achieve what a clear comparison and one controlled action could have completed in two.
These signs do not prove that A2UI is the answer. They tell the product team where to investigate: the journey between the user's need, the information presented and the next useful action.
A bad experience does not stay in the interface
The customer experiences effort. The business experiences the consequence.
Some people will ask the agent another question. Others will leave, contact support, copy the answer into another tool or avoid the product next time. They may never report that the interface was the problem. The business simply sees an incomplete purchase, a longer handling time, another correction or an AI feature that appears to have weak adoption.
| What the user experiences | What the business may experience |
|---|---|
| “I cannot see the difference between these options” | A delayed decision or an abandoned purchase |
| “I have already given the agent this information” | Repeated work, longer handling time and higher support demand |
| “I do not know whether the change has happened” | Duplicate actions, corrections and reduced trust |
| “I will copy this somewhere easier to use” | Broken workflows, missing context and avoidable data errors |
| “This tool makes my job harder” | Low repeat usage and an AI investment that appears to fail |
A technically correct answer can still produce a poor business result. If the useful number is hidden in paragraph four, the customer may miss it. If the next action is unclear, they may stop. If an action happens without a visible review step, one mistake can damage confidence in the whole product.
This is why message count is a weak success measure by itself. Six messages might represent a useful conversation. They might also represent five failed attempts to complete a simple task. Measure what happened after the conversation: decisions made, work completed, corrections required, support requested and whether people chose to return.
How the chat box became the default
The original ChatGPT did more than popularise a model. It gave a large number of people a mental model for using AI: write what you want, receive a response, then continue the conversation.
That was a sensible beginning. Natural language is forgiving and remarkably versatile. A button needs a label and a known action. A conversation can begin before either side fully understands the problem.
That versatility is an advantage in a general assistant because the next request could be almost anything. It can become a disadvantage as we build agents for narrower jobs. If an agent is designed specifically to qualify a sales lead, compare insurance options or approve an expense, asking the user to describe every known input in a blank box creates unnecessary freedom. The user must work out what to say, the agent must interpret it and the business must handle more ways for the request to be misunderstood.
A specialised agent should still make room for an unusual situation. But it should not pretend that every routine step is an open-ended conversation. The more clearly the job is defined, the more useful visible choices, status, evidence and actions become.
But the model's capabilities expanded faster than that interaction pattern.
In March 2023, ChatGPT plugins began connecting the conversation to current information, calculations and outside services. The model was no longer limited to composing a response from what it had learned during training. It could begin to retrieve and do things. Read the original plugins announcement.
In October 2024, OpenAI introduced Canvas for writing and coding work that went “beyond simple chat.” A user could select part of a document, edit it directly, receive inline suggestions and restore an earlier version. OpenAI described Canvas as the first major change to ChatGPT's visual interface since the 2022 launch. Read the Canvas announcement.
In October 2025, apps in ChatGPT combined conversation with interactive interfaces such as maps, playlists and presentations. The conversation could remain the starting point without requiring the answer to remain a message. Read the apps announcement.
By then the direction was becoming clearer. When AI helps somebody edit a document, inspect a map, compare products or complete a workflow, a sequence of messages is not always enough.
Chat opened the door. The work kept asking for more.
Ask in ordinary language and continue with follow-up questions.
The conversation begins reaching current information and outside services.
Writing and coding gain a workspace beside the conversation.
Maps, choices and other useful controls appear inside chat.
The agent can describe a task-shaped interface for an approved renderer.
Chat did not become obsolete. It became one part of a larger interface.
A product that gets this right
One experience I keep returning to is Gemini inside Google Docs.
If I ask an ordinary chatbot to improve a paragraph, it often returns another paragraph. I then compare the two versions in my head, decide what changed, copy the new text and hope I did not replace something I meant to keep.
Gemini in Docs can place suggested edits where the work already lives. The difference is visible in context. I can review one change, accept it, reject it or accept all changes. The AI proposes; I remain the editor.
That sounds like a small interface decision. It removes four pieces of work:
- no copying text into a separate chat;
- no holding the old and new versions in memory;
- no hunting through a long answer to find the proposed wording; and
- no ambiguity about whether the document changed before I approved it.
Watch the 27-second demo ↗Gemini's “Help me write” experience inside Google Docs. The 27-second demo shows a proposed edit in context, the visible difference and the controls for accepting or rejecting it. This screen capture was taken on 21 September 2026 from a video Google published with its 10 March 2026 product announcement. It represents that release, not every account or the current product. Video © Google. Use the preview above or
open the official Google Workspace video in a new tab.
The accompanying announcement is on the
Google Workspace Blog.
Google's current help documentation describes the same control model: suggestions appear directly in the document, and the user can accept one, accept all or reject all. See how suggested edits are reviewed in Docs.
This is the point in miniature: the prompt starts the work, but a purpose-built interface makes the result easier to judge and safer to apply. A business does not need to copy Google Docs. It needs to notice the principle. Put the proposed change beside the thing it affects. Make the difference visible. Let the person decide.
Text and controls are good at different things
Return to the plan selection.
The user may begin with a sentence because their need is still untidy:
We have 12 people now, may hire four more this year and need faster help when customer enquiries get stuck.
That is good work for language. A fixed form might ask for team size and support level but miss the expected growth and the reason faster support matters.
Once the choices are known, the design principle changes. It is easier to recognise a relevant option on screen than to remember that it exists and type its name correctly. Good interfaces keep important choices visible instead of asking the user to carry them between messages. See Nielsen Norman Group's explanation of recognition and recall.
Once the agent understands the situation, known choices can become controls:
- a team-size field already set to 12;
- a monthly or annual billing choice;
- a live-support requirement;
- two plan cards that meet the basic conditions;
- a feature-comparison table;
- the price difference in the user's currency; and
- a visible action to select a plan or ask another question.
The person can still type, “What happens if we grow to 25 people in six months?” Conversation handles the exception. The controls handle the ordinary comparison.
A capable answer can still make the user do too much work.
Which plan should I choose for 12 people?
The Starter plan supports 5 people for KES 4,500 monthly. Team supports 20 for KES 12,000 monthly with email and chat support. Business supports 50 for KES 28,000 and includes priority support, but is billed annually…
Which ones have live chat and monthly billing?
Team has live chat and monthly billing. Business has priority support but requires annual billing. Based on your team size, Team may be the most suitable option…
- Monthly billing
- Email and live chat support
- Room for 8 more team members
Business also fits, but requires annual billing.
This is the sweet spot. It is not a compromise where half the screen must be text and half must be components. It is a division of labour.
| The user needs to… | Useful first interaction |
|---|---|
| Describe an uncertain or unusual need | Natural-language text |
| Choose from known valid options | Select, toggle, radio buttons or plan cards |
| Compare exact features or prices | Table with consistent rows and columns |
| See change, scale or an unusual pattern | Appropriate chart with labels and source context |
| Understand a recommendation | Short explanation with evidence and uncertainty |
| Complete a consequential action | Explicit review and confirmation control |
| Ask about something the designer missed | Conversation |
The user should not have to choose one interface philosophy before beginning. The experience can change shape as the work becomes clearer.
Use language for intent. Use the interface for the work.
A chart should reveal something
Visual presentation is not a cosmetic reward for completing the analysis. It affects what people can notice.
Suppose the user asks which branch changed most this quarter. A paragraph can report every percentage correctly and still make the comparison tiring. A well-labelled bar or line chart can make the outlier visible before the user finishes reading the heading.
Research into graphical perception has long found that people compare some visual encodings more accurately than others. Position along a common scale and length generally support more accurate comparisons than angle or area. That is one reason a straightforward bar chart will often serve the reader better than a dramatic collection of circles. Read Cleveland and McGill's foundational graphical-perception paper.
But “make it visual” is not a requirement to make everything a chart.
Use a chart when the reader needs to see:
- change over time;
- the difference between categories;
- progress towards a target;
- concentration or distribution; or
- an unusual value that deserves investigation.
Use a number, sentence or table when that communicates the answer faster. A chart should reveal a pattern, not provide colourful accommodation for three numbers that were perfectly comfortable in a sentence.
The visual also needs labels, a clear scale, sufficient contrast and a text or table alternative. If a chart makes the answer faster for one person and inaccessible to another, the design has moved the work rather than reducing it.
Where A2UI enters the journey
So far, we have described the customer problem. A2UI is one emerging engineering approach for delivering the better experience.
A2UI means Agent to User Interface. It is an open protocol, initially created by Google and developed with open-source contributors, that allows an agent to send a structured description of an interface to another application. That application renders the description using components it knows and trusts. Visit the A2UI project.
In plain language, the agent can request a plan comparison, two eligible options and a confirmation action. The company's application decides how those pieces look and what they are allowed to do.
The current A2UI 1.0 specification is marked as a candidate and remains an actively developing standard. It describes streamed interface updates, a separate data model and component catalogues that define which interface elements are available. Read the A2UI 1.0 candidate specification.
You do not need to remember those technical details. The useful distinction is this: the agent can arrange approved building blocks for the task without being allowed to invent arbitrary behaviour.
What management should insist on
An adaptive interface still needs firm boundaries. For the plan-selection journey, I would ask for five things.
Approved building blocks
The agent may use the company's plan cards, comparison tables, inputs, notices and confirmation actions. It should not be able to create an unexpected control because the user sounds impatient.
Visible uncertainty
If the customer says reporting matters but has not explained which reports, the interface should show Needs confirmation. It should not quietly label the cheapest plan Recommended.
Real permission checks
A visible Choose Team button does not prove the user may change the subscription. The application must verify identity, authority, current pricing and required billing information outside the generated interface.
A clear consequence
Before confirmation, show the exact cost, billing period, renewal terms and what will change. A polished button can create more confidence than a paragraph, which makes a misleading button more dangerous.
A safe fallback
If a component cannot load or an action fails, the customer should still understand the answer and know whether anything changed.
These controls matter whether the product uses A2UI, another generative-interface approach or a conventional screen designed by hand. The protocol can carry an interface description. It cannot decide what the business should permit.
A practical review for an AI product
Choose one task people begin but do not complete. Watch how they move through it, then ask:
- Beginning: Do customers know what they can ask or do?
- Input: Are they typing known information that a control could collect faster?
- Memory: Must they remember choices that could remain visible?
- Response: Is the answer obvious, or hiding in paragraph four?
- Comparison: Would aligned rows, cards or a chart reveal the difference faster?
- Action: Is the next permitted step visible and specific?
- Trust: Can customers see missing information, important assumptions and the consequence of acting?
- Completion: Are you measuring finished work or merely conversations?
Useful measures include time to a decision, tasks completed, questions repeated, invalid submissions prevented, abandoned journeys, corrections, reversals and whether people voluntarily use the workflow again.
Start with one journey, not a universal interface
A2UI is still emerging. I would not begin by asking an agent to redesign an entire business application differently for every user.
Start with one bounded journey where text is visibly doing too much. Plan selection is a good example because the options are known, comparison matters and the final action can remain controlled.
Test the experience with a small group. Let people describe the need naturally, collect known parameters with controls, keep important differences aligned and preserve conversation for unusual questions. Then compare completion time, corrections and abandonment with the chat-only journey.
The protocol is not the value by itself. The value is helping a person understand and complete the work with less effort and fewer avoidable mistakes.
The chat box was the beginning, not the mistake
The original ChatGPT interface succeeded because it made a powerful new capability approachable. It gave people permission to begin without understanding the machinery underneath.
We should keep that lesson. People should still be able to arrive with an incomplete thought and say what they are trying to do.
But once the task becomes clear, the interface can help carry the cognitive load. It can keep choices visible, align the comparison, reveal the pattern, collect valid information and show the next safe action.
My view is that the future is not chat or graphical interfaces. It is a thoughtful movement between them:
Use conversation to understand the person. Use structure to organise the answer. Use visuals to reveal what matters. Use controls to make the next step clear.
A2UI is an early attempt to give agents and applications a shared language for that movement. Whether this particular protocol becomes widely adopted is still an open question. The underlying product lesson is already useful.
As agents become more powerful, we should not require people to become more skilled at extracting value from paragraphs. A better system should know when to stop talking and help the person see what to do.
For a customer-facing example, follow the AI and WhatsApp order journey from product search to cart approval and an order in the business's existing system. Continue with how AI integration changes management information for the connection between conversation, business records, visuals and controlled action. If the interface may trigger changes in an ERP, CRM or payment system, read how to define an AI agent's permissions.
Research and helpful links
Introducing ChatGPT in 2022
Early ChatGPT interface reproduced in the AASA Journal of Scholarship and Practice
Introducing Canvas
Introducing apps in ChatGPT
Google Workspace's 10 March 2026 Gemini product announcement
Google Workspace's Help me write demonstration
Google Docs guidance for reviewing Gemini's suggested edits
A2UI project overview
A2UI 1.0 candidate specification
Recognition and recall in user interfaces
Graphical perception research by Cleveland and McGill
