Chat is the wrong interface for AI
The problem with over-simplification of AI Agents’ interfaces.
Chat is the wrong interface for AI
The problem with over-simplification of AI Agents' interfaces.
I noticed that I'm becoming increasingly annoyed while talking to AI agents. Not because they are not intelligent enough, but quite the opposite.
I can spend an hour discussing an idea with a model, while walking with my dogs. We're moving between design, engineering, psychology, physics, and a real product that appeared somewhere in the middle. The LLM can follow me surprisingly well.
It can recognize that two things I said at different moments are connected, follow my return to an earlier hypothesis, challenge me if I ask it hard enough, and sometimes formulate a thought I was circling around but could not express clearly.
And all of this intelligence is trapped inside a messenger, like a chatbot from 2015 but on steroids.
But the conversation becomes longer, the scrollbar becomes smaller, and my best ideas disappear somewhere between 10 tool calls, 3 deliverables, and 100 affirmations that my question is fascinating and my idea is a goldmine.
As a designer, I see an interface problem. But as a person who studied linguistics, psychology, and communication theory before switching to design, I cannot look at it only as a usability problem.
And since we are talking about Large Language Models, and language is still the primary interface for accessing the technology, the very object of today's discussion will be the meaning of language, the construction of thought and the strange decision to place some of the most advanced computer systems ever created inside an interface inherited from SMS and internet chat rooms.

We can do better.
T
he model of language
A modern language model can process enormous amounts of text, identify patterns across distant parts of a conversation, synthesize information from many sources, produce software, analyze images, and operate tools. But from the user's perspective, all of this is normally reduced to one interaction:
I write > Machine answers > I write > Machine answers
The action-reaction structure is so familiar that we rarely question it. Chat feels natural because conversation has been the main way we humans communicate for thousands of years. It's communication theory 101.
But natural does not mean complete.
Chat presents communication as a chronological record of messages. It's excellent at showing the order in which something was said. But it doesn't show the incremental change of state, which ideas are now connected, which assumptions were rejected, what remains unclear, and what the conclusions are after 30 minutes and 6 iterations of debate.
I would describe it this way: Chat is not the structure of communication. Chat is its serialization.
Why language needs time
At the beginning of the 20th century, Swiss linguist Ferdinand de Saussure described the "linear nature of the signifier" in Course on general linguistics. An auditory sign unfolds in time, which means that we can't use every word, relation, nuance and reference simultaneously. One sound, morpheme, word follows another. The listener must reconstruct the meaning while the sequence is still arriving.

And another reason to speak linearly is that only one person can comfortably occupy the channel of verbal communication at a particular moment. Thus, we, as a social species, have developed a highly organized system around conversations.
In a paper A Simplest Systematics for the Organization of Turn-Taking for Conversation, Harvey Sacks, Emanuel Schegloff, and Gail Jefferson show that ordinary conversation is not a random exchange of noises. We constantly predict when another person may finish, interrupt, pause, repair misunderstandings and constantly adapt our message to the reaction of our opponent.
This is one of humanity's most impressive but invisible interfaces.
We do not need a loading indicator; we hear in the pause that another person is searching for a word. We don't need a system notification to understand that a sentence is not finished, because we use intonation. And we can start forming our answer before the thought has been fully expressed because spoken-language processing is incremental and predictive, as stated in research on prediction during language processing by Susanne Brouwer and her colleagues.
So linearity is not that primitive or a mistake that we should simply remove. It's a solution to the physical reality of communication in time.
But... the sequence in which information was delivered is not necessarily the structure in which it remains inside of our cognition.
Language is not a thought
This distinction feels intuitive, but it also has a serious scientific foundation.
In 2024, an article with the direct title Language is primarily a tool for communication rather than thought by E. Fedorenko, S. Piantadosi and E. Gibson was published in Nature. Based on neuroscience, they argue that language and thought can be dissociated. Language is very important for transmitting knowledge, but it doesn't appear to be a prerequisite for different forms of cognition, including symbolic thought.
This does not mean that language plays no role in thinking; of course it does. We experience inner speech, rehearse memories and make abstract ideas easier to manipulate. Language can participate in thought without being identical to it.
A linear sentence can describe a spatial system. A sequence of words can express two conflicting possibilities. A chronological story can contain memories, predictions, parallel events and references to something that happened three chapters earlier.
The line carries the model. It's not the model itself.
We learn to begin somewhere, continue somewhere else, and eventually stop. Not necessarily because the thought arrived in that order, but because another human being cannot enter all of it at once.
Without this discipline, we would drown each other in chaos.
But what if the machine could take over this "translational" part? Not the work of thinking for us, and not the work of deciding what we mean, but the exhausting work of continuously reorganizing what has already been expressed into a structure that all participants can inspect.
Imagine a software repository. The commit history shows us how the product changed: what was added, removed, broken, repaired, and reconsidered. But when a developer wants to understand the system as it exists now, they do not reconstruct the entire application by reading every commit from the beginning. They inspect the current codebase.

Our conversations with AI work in the opposite way. We keep the commit history permanently open and ask both ourselves and the model to reconstruct the current state from it every time.
No wonder the context gets lost.
So what shape does a thought have?
This is where we need to be careful. It's tempting to say that the human brain thinks in graphs. It is a beautiful metaphor, especially for a designer's mind, because we immediately imagine nodes, connections, clusters, and the much-loved infinite canvas. But a visual graph is not a biologically accurate picture of thought, and I would rather not replace one simplification with another.

What we can say is that cognition often behaves in ways that a simple line represents poorly.
In Spreading-Activation Theory of Semantic Processing, Allan Collins and Elizabeth Loftus proposed that activating one concept makes related concepts more accessible. Other neuroscience research has described semantic memory not as one isolated storage area, but as a large distributed system involving different kinds of information across the brain, as Jeffrey Binder and Rutvik Desai explain in The Neurobiology of Semantic Memory.
When I'm thinking about my dog, I may activate a visual image, a sound of barking, the weight of a leash in my hand, a memory of walking in the rain, the concept of loyalty, and the knowledge that I still need to buy food... These elements do not politely wait for their turn in one semantic queue. They appear through association.

Creative thinking makes this even more visible. Studies of semantic networks suggest that highly creative people tend to have more flexible associative structures, allowing movement between more distant concepts. Research by Yoed Kennet and colleagues found differences in the organization of semantic networks between people with lower and higher creative ability, in Investigating the Structure of Semantic Networks in Low and High Creative Persons, and later explored this flexibility computationally in Flexibility of Thought in Highly Creative Individuals Represented by Percolation Analysis.
This sounds familiar to anyone who has participated in a workshop. One person places an idea on the board. Another person connects it to a user problem. Someone remembers an earlier research insight; a contradiction appears. Two clusters become one. The original question is reformulated, and suddenly the team sees that it has been solving the wrong problem for the last hour, but the meeting still happens in time.
Philips Johnson-Laird's work on mental models gives us another useful layer. People construct internal representations of situations and use them to consider possibilities and alternative realities. This doesn't mean we can directly display someone's private mental model on a screen. But we can build an external working model from a person's input and allow them to correct it.
So, the AI should propose the structure, but it's essential that the user owns the meaning.
We already think outside our heads
As designers, creators, builders, we do this all the time. We cover walls with sticky notes, create diagrams, service blueprints, CJMs, gigantic Miro boards that begin as a workshop and eventually become a wasteland no one wants to visit again.
We call these things documentation or visualization, but that is only partly true.
Moving a sticky note is sometimes an act of thinking. Placing two ideas next to each other allows us to compare them without keeping both fully active in working memory. Drawing a line can reveal a dependency we did not see before. Separating an assumption from evidence can change the conclusion.

This visual space does not only show the result of cognition; in fact, it participates in the cognitive process.
David Kirsh and Paul Magilo called similar operations epistemic actions - actions that are performed not primarily to change the external world, but to make a cognitive task easier or less error-prone. In their example, skilled Tetris players rotated pieces on the screen instead of completing every rotation mentally. The action changes the environment so the environment can help calculate the answer, as we delegate the building of a mental model to the physical world and observe its results before taking the next action.
And this is exactly what I want from an interface with an AI agent. I don't want the machine only to give me an answer that I then have to keep somewhere inside my head. I want the environment between us to become part of the thinking system. I want to see the assumptions we are currently using, the questions we have not answered, the evidence that changed our direction, and the ideas that may be irrelevant now but should not be forgotten forever.
The difference between text and a diagram is not simply visual preference. In Why a Diagram is (Sometimes) Worth Ten Thousand Words, Jill Larkin and Herbert Simon showed that diagrams can group related information spatially, reduce search, and support perceptual inferences that are difficult to perform through verbal interpretations.
The word "sometimes" in the title is very important. Not every diagram is better. We have all seen graphs that look like a plate of spaghetti. A bad spatial representation can hide meaning just as effectively as an endless chat history.
So the answer is not one enormous mind map. The answer is an adaptive system of different forms of mental-model representation.
The problem with chat
Let's get back to our initial topic. Chat is excellent when I need a local answer. I ask how to change a configuration, translate a paragraph, understand one term, or compare two products, and it gets the job done. The action-reaction format is quick, familiar, and often exactly enough.
But the moment I dive into real, continuous work on a product, the complexity of context and thought process exposes its limits.
If I open several branches of thought, they mix inside the same timeline. If I return to an earlier assumption, the interface does not clearly show which later conclusions depended on it. If the agent misunderstands my goal and we correct it ten messages later, the wrong version remains visually equal to the right one. A confirmed decision, a temporary hypothesis, and an unresolved question all look like the same object: a message bubble, in the best-case scenario, some artifact created by an LLM.
The interface displays syntax (a word order) where I need semantics (a meaning).
HCI researchers have already begun documenting this problem. During the study behind Graphologue, participants found LLM responses verbose and time-consuming to comprehend, while the linear structure created repetitive scrolling, copying, and pasting during complex information tasks. Graphologue responded by transforming model output into interactive node-link diagrams.

Sensecape went further by allowing users to move between a spatial canvas and a hierarchy, supporting exploration at different levels of abstraction. In its study, users explored more topics and organized their knowledge more hierarchically than through a conventional interface.

And Branchat experimented with a tree-structured conversation where users could revisit an earlier moment, create a new branch, and control which context should remain active. Basically trying to apply git a like approach but for conversations.

These systems and ideas do not prove that chat is dead; they give a perspective on an alternative view for managing communication's context.
The first cracks in the wall
What fascinated me most during this research was discovering how similar skepticism led to different ideas and approaches, and how close some experimental structures are to the interfaces I had in my head.
And besides Graphologue, Sensecape, and Branchat, which we've discussed in the context of structures, there are other projects that are also thinking in the same direction. These projects are interesting not only because they are also trying to build with research in mind, but also because of the vision and the use they are trying to find for this alternative informational structure.
MeetMap uses a large language model to generate collaborative dialogue maps during online meetings. It contains both a chronological topic timeline and a map canvas. The timeline explains how the conversation developed and the map helps participants organize what the conversation currently means.

History and state are not competitors, as I said before. We need both, to know how we arrived here and to understand where we are now.
Just recently, Orality: A Semantic Canvas for Externalizing and Clarifying Thoughts with Speech was published, a paper on the system that extracts key information from spoken language, by using an LLM to build a node-link diagram. The system allows users to manipulate clusters, accepts verbal commands to reorganize the content, proactively proposes new questions and detects conflicts. In a study, Orality supported the clarification and development of thoughts better than voice interaction with ChatGPT.

This is already the next generation of transcription and thought-compression tools.
MindTrellis explores co-creation with AI. Not AI generating finished knowledge for the user, but the user and AI collaboratively maintaining a knowledge graph, where people can introduce concepts, modify relationships, reorganize hierarchies, and query the evolving structure.

So the researchers are already making cracks in the wall of chat. But all that I've mentioned above still focuses on a particular session, task or meeting. They visualize one conversation or help organize one information space. But it's rarely how we humans operate.
The larger opportunity begins when this structure stops being owned by one chat or output and becomes the persistent context in which agents and people operate. Just one abstraction above.
Memory is the first practical step
Before we redesign chat-structured communication, we probably need to redesign memory.
Current agents often remember by placing earlier information back into the model's context: raw chat history, retrieved fragments, summaries, user preferences, or some combination of them. This is useful, but it still treats memory largely as text that must be selected and injected into another text sequence.
Recently, Andrej Karpathy proposed the LLM Wiki pattern. Instead of forcing an agent to rediscover the same knowledge in every session, it incrementally compiles raw sources into a durable and human-readable wiki. The sources remain available, while the agent maintains a synthesized knowledge layer that can be revised, queried and expanded. All as a simple file-/folder-structure.
Similarly, Microsoft Research's GraphRAG uses an LLM to extract entities and relationships from a large collection of documents, organize them into communities and create layered summaries, in a more complex system than LLM Wiki.

These approaches help machines compress information into models they can navigate. But then we do something strange. We organize the information relationally for the machine and flatten it back into chat for the human.
The better approach would be to bring this mechanism into live communication itself.
Imagine an LLM Wiki that does not begin only with documents placed into a folder, but with the language continuously produced between a person and an agent. Not every sentence would become a permanent memory. But useful concepts could move through different layers: immediate context, working context, episodic memory, stable semantic knowledge, and eventually an archive or so-called long-term memory.
The agent would not merely consume the wiki; we would watch it being built, we could correct it and disagree with it, and sometimes we could tell it to forget or it would even "forget" by itself, which means adapting. The same way we adapt our nonlinear thoughts in linear communication.
From message history to shared context
The central object of such an interface would no longer be the message or the node itself, but the context.
Not context in the common understanding of a RAG system or LLM Wiki, but the omni-present one, with the current state being built in real time as well as all previous branches of it being present.
Chat could remain as one way of interacting with it, especially when conversation is the simplest and most human method. But besides the chat, there could be a living model containing concepts, claims, goals, questions, evidence, assumptions, decisions and relationships ready to adapt and transform into an alternative interface: a kanban board, mindmap, timeline, diagram, or something more interactive.
Not all of them should look or behave the same.
A statement directly expressed by the user is different from an interpretation made by the AI. A temporary working hypothesis is different from a fact supported by 3 sources. And a conclusion accepted by one person in a meeting should not be visualized as common agreement among everyone.
Research on grounding in communication by Herbert Clark and Susan Brennan gives us an important foundation. Shared understanding is not created because something was said. Participants continuously establish evidence that information has been understood well enough for their current purpose.
So a shared context system would need states, not only nodes: expressed, inferred, questioned, accepted for now, supported, unknown, etc.
And instead of only giving us another answer the agent could show a context diff (what changed in our shared model as a result of the interaction).

This would allow us to understand not only what the agent said, but what it now believes is relevant, which part of my model it thinks has changed and where our interpretations may differ.
We could even visualize several models at once:
- the model I expressed
- the model the AI inferred
- the part we accepted as shared
- the version we had yesterday
- the version of another participant
- etc.
Misunderstanding would not disappear, of course, but it could become visible earlier.
Not a mind map, but a context interface
I love mind maps, but I don't think a mind map alone is the one and only future of AI interfaces. As I wrote above, I believe that this future is adaptive and can have different faces. From graph to timeline, and from table to document.
The deeper idea is not to replace one universal interface with another. It is to maintain one underlying context that can be expressed through different views.
This is why I see it almost as a Context Operating System: a layer above individual agents and models, skills and plugins. It should be another level of abstraction: a system that doesn't belong to a session, harness, model, or company, but rather its own instance, stable and flexible at the same time.
Ontology... what?
Until this point, I have deliberately avoided using the term ontology, even though this seemingly philosophical but in reality computer science concept is appearing more and more often in discussions around artificial intelligence and memory management.
I'm sure that many of you already had this idea in mind and were wondering why I had not mentioned it yet. Introducing this category would have taken the complexity of this article to an entirely new level and pulled us away from its main subject.
And honestly, this topic deserves an article of its own.
But I also don't want to leave the question completely unanswered. I recommend watching a short lecture by Frank Coyle, Why Agentic Systems Need Ontologies. In it, he suggests treating reasoning as an inherently probabilistic process while supporting it with logical representation through older but stable standards such as RDFS and OWL, and of course graphs.
Before the neuro-interface
It's easier than ever for us to imagine the future of communication through brain-computer interfaces. Thoughts transmitted directly, language becoming unnecessary, and meaning moving between people without the limitations of speech.

Perhaps one day.
But before we connect a machine directly to the brain, we can learn to build interfaces that operate one level above sentences.
I would call this a form of meta-linguistic communication. Language remains the transport layer because it is flexible, deeply human and remarkably effective. But instead of consuming only the linguistic output, the system maintains a model of the context that the language is continuously editing.
And we would be in the same workshop that we started a few thousand words ago but this time we would share an interface that not only transcribes the conversations but continuously displays emerging agreements, competing interpretations, unsupported assumptions and unanswered questions. Imagine being able to ask not "What did Maxim say twenty minutes ago?" but "How does Maxim's current model differ from mine, and where did that difference begin?"
This is not mind reading. It is better tooling for the difficult work humans already do when we try to understand one another.
Conclusion
I'm more than certain that major changes are coming, and we need to prepare for them. The chat interface was a great starting point. Even 2 or 3 years ago, it would have been difficult for me to imagine anything else taking its place.
But progress, the speed of that progress, and the constant stream of information we face every day push me to question the current state of interfaces for AI.
It genuinely hurts me to begin a new large project or task by creating yet another chat whose context window will be overloaded after just 20 minutes of discussion. The model gradually becomes less capable until it reaches a boiling point and has to compress the context through a process that is completely opaque and beyond my control, and that will inevitably result in some information being lost.
This is why I experiment. I'm trying to understand the very essence of the problem.
I create alternative interfaces and boards (that your coding agent can build from a single prompt) helping me reduce my mental load and delegate tasks to agents more effectively. Kanban and ticketing systems are familiar to me, so they give me a structure I already understand.

I play with context and with the idea that, instead of creating one context for an entire session, we could provide an LLM with a separate context for every individual prompt, and manage all of it within a node-based system.

I experiment with different ideas and input methods, starting with physical controllers and moving toward interaction through sign language or eye tracking, with the goal to replace or complement text input, whether typed or spoken.

And, of course, if I already had the answers to the questions I raise in this text, this material would not exist. There would be a new page on ProductHunt instead.
That is exactly why I'm writing here, where curious and thoughtful minds come together. I hope these questions will interest someone else, encourage them to begin experimenting too, and help us bring our bright future a little closer.
And if you are interested in this topic, here is some further reading and viewing from authors who are also exploring AI interface design and nonlinear thinking:
- Where should AI sit in your UI? by Sharang Sharma - an overview of emerging UI patterns and how they shape the AI experience.
- The UX of drafting in space by Daniel Buschek - a practical exploration of nonlinear thinking through the experience of working with Miro.
- Imagining better interfaces to language models by Linus Lee - written near the beginning of the current AI boom, yet, in my opinion, it has lost none of its relevance. It raises fundamental questions and remains a great source of inspiration.
- Why Agentic Systems Need Ontologies by Frank Coyle - the video I mentioned earlier in the article. It addresses some of the same questions raised in this text, but from a more technical perspective.
Source: Medium

Exclusive LaunchPad
30% Off
