Request a demo
AI

The questions enterprise assistants can't answer

· Ed Sherrington
The questions enterprise assistants can't answer

Most teams building an internal assistant end up in one of two camps, usually within six months of starting.

The first views the core challenge as a retrieval problem. Convert everything the organisation has written into markdown, chunk it, embed it, and let the model find what looks relevant. Policies, wikis, decision records, Confluence pages, the deck someone made for the steering committee. It is quick to build, covers an enormous surface area, and demos beautifully.

The second treats it as a modelling problem. Define the entities that matter, the relationships between them and the rules that govern them. Build a formal ontology, put a query layer over it, and let the model translate questions into queries against something known to be true. It may be slower to do, but every answer it produces can be defended.

Each camp has weaknesses. Documents are loose, stale and hard to reason over. Ontologies are narrow, expensive and perpetually behind.

Yet both arrive at roughly the same place in production. They handle the questions people asked during the pilot, and stall on the real-world questions that would have justified the spend.

Where documents fail

This failure is the better understood of the two. The key points are:

  • Similarity is not relevance. Retrieval finds text that resembles the question. Resemblance and truth are different properties, which is why models answer incorrectly with complete confidence.
  • Anything needing more than one hop degrades. Per the MultiHop-RAG benchmark (Tang and Yang, COLM 2024): GPT-4 answered 56% of multi-hop queries correctly from retrieved chunks, against 89% when handed the ground-truth evidence directly.
  • Documents describe a moment. Nobody updates documents when their subject matter changes. Retrieval cannot tell which parts are still true.

Where ontologies fail

For a bounded domain with known interests, a formal model is hard to beat. Instrument reference data, a regulatory reporting hierarchy, a product taxonomy: stable meanings, a known question set, a high cost of being wrong. Here, ontologies make sense.

The trouble starts when the same method is applied to the institution as a whole. Because an ontology’s schema is a set of subjective decisions about which distinctions matter, taken before anyone has asked questions. Four things follow:

  1. The schema has to anticipate the question. In a bounded domain that is achievable. Across an institution it isn’t. A question the model was not built for isn’t a slow query, it’s an unanswerable one. Fixing it means a full modelling exercise, a design review and a change board.
  2. The valuable questions cross the boundaries the schema drew. Ontologies get built by the teams that own the domain: payments models payments, risk models risk, and so on. The questions worth asking run between teams, with no single owner positioned to build the thing that spans them.
  3. Most ontologies describe the present. They hold current state well and history poorly. “What is the position” is answerable. “How did we get here, what did it depend on, when did it start drifting, who was involved” is not, and that second set is where the value often sits.
  4. They model the organisation as designed, not as it runs. Formal models are tidy, institutions are not. The same supplier may be under four different names across three systems. Programmes that are one line in a portfolio report may be three teams in practice. And of course there are plenty of examples of “that’s just how it is” knowledge that everybody works around but nobody has ever formally recorded.

None of this can be fixed by building a bigger ontology. An ontology’s rigidity isn’t a defect of its schema, it’s the property that makes its schema useful. Every modelling decision buys determinism in one direction, at the expense of foreclosing a class of questions in another.

The useful questions are the ones nobody modelled

Consider the kind of request that is actually useful to a senior leader: “I’m going to London to meet X. Tell me what I should know about them, and what we’ve done with them recently.”

On the face of it, this is not a complex or exotic question. But answering it crosses everything. Coverage, transactions, complaints, meetings, the escalation in March, the pricing exception somebody granted last year.

An ontology can rarely answer these sorts of questions because they don’t belong to a single domain and thus weren’t modelled ahead of time. A document corpus can have a go - but what they assemble is rarely reliable because the answers don’t sit neatly in specific documents.

Time as the fundamental gap

Look again at that question’s phrasing: what we’ve done with them recently. This might mean “since we last spoke”, “since I last read about them”, or “in the last week”. Almost every question worth asking carries a clause like this, and neither school of thought is equipped to deal with it.

A document corpus has timestamps on files, not on facts. It shows when a page was written. It cannot show when something became true or when it stopped being true, which is why a retrieved paragraph is so often accurate about a world that no longer exists.

An ontology holds the present, and stays useful by overwriting. The new owner replaces the old one, the status moves from amber to green, the dependency gets re-pointed. Each update keeps the model current by destroying the record that would have let anyone ask how it came to be that way.

So the institution ends up with one system that remembers everything and cannot say what is still true, and another that knows exactly what is true today and nothing about yesterday.

Structure and language, running together

Neurosymbolic systems offer a third path for enterprise AI. They process structure and language together, rather than choosing one over the other.

The symbolic half holds entities and relationships. It supports deterministic traversal and exact match, and it can show the path that produced an answer. The neural half handles language and everything the structure does not cover. It maps a question onto a model nobody phrased the same way. It reads the prose that never got formalised. It recognises that four supplier names refer to one supplier.

More advanced neurosymbolic systems hold time as part of the symbolic structure (rather than as metadata attached to it), so every entity and relationship carries the period over which it held. This means history is preserved - without compromising on performance - so models can reason across structure and time, not just structure.

An important point with these systems is that both sides run in parallel, not in sequence. Structure and meaning are searched together and the results compose. This is different to building an ontology with vector search as a fallback. The fallback option routes questions between two deficient systems. Parallel processing enables the whole to be greater than the sum of its parts.

What to use, where

Neurosymbolism trades well in multiple directions: trust comes from the symbolic side, coverage comes from the neural side, and complexity is handled by combining both. But it is not a silver bullet to solve all enterprise AI problems.

Structure still has to be committed to. The approach doesn’t remove modelling, it just lowers the cost of getting the model wrong (because the neural layer can reach what the schema missed). And its reliance on generative AI means answers are still probabilistic: models can select the wrong path with total confidence. Provenance doesn’t prevent that - it just means answers are checkable and mistakes, visible.

So for high stakes questions about a tightly-bounded, known domain: a formal ontology may well be the right path. And for simple, low stakes questions about well documented and up-to-date repositories: documents may still be the right solution.

But for institutions serious about enterprise assistants that operate effectively in the wild, answering both the questions that weren’t anticipated and the complex ones that were, a neurosymbolic approach is the way forward.

Institutions are not tidy and they do not hold still. Any assistant that treats them as if they are may demo well, but will leave the important questions unanswered.

If you want to talk about what a neurosymbolic approach would look like against your own systems, get in touch.

Experience Pometry in action
with a live demo.

Find out how you can close your visibility gap with unique, institutional intelligence.