Why the quality, visibility and governance of enterprise data increasingly determine whether generative AI delivers value

August 24, 2026

Adrian Knapp is founder and CEO of APARAVI (Source: APARAVI)

Generative AI has become a strategic priority for companies across almost every industry. The debate usually focuses on more powerful large language models, new agentic systems or the return on investment of AI projects. Far less attention is paid to the foundation on which every successful AI application depends: the underlying data – and, in particular, unstructured data.

APARAVI, a provider of technology for preparing unstructured enterprise data for AI use, argues that many organisations are looking for the causes of failed AI projects in the wrong place. When results disappoint, the language model is often blamed, prompts are considered too imprecise or AI agents are accused of working unreliably. In many cases, however, the problem lies much deeper.

Most enterprises still have a far better understanding of their structured data than of the information contained in emails, documents, meeting minutes, videos, presentations, collaboration platforms and file servers. Yet this unstructured information often contains a substantial share of the knowledge required for operational and strategic decisions.

If an AI system can access only fragments of that knowledge, even a highly capable model may produce answers that sound convincing while being based on an incomplete picture. APARAVI identifies six areas that organisations should address before scaling generative AI.

1. AI often sees only a fraction of corporate knowledge

Companies have invested heavily in data warehouses, business intelligence systems and analytics platforms. These environments are designed primarily for structured information such as transactions, tables, KPIs and numerical datasets.

A large proportion of organisational knowledge, however, remains outside these systems. Project documentation, emails, technical specifications, meeting notes, presentations and collaboration platforms frequently contain the contextual information needed to understand why decisions were made, how processes evolved or what risks exist.

If generative AI can access only structured repositories, it may provide technically plausible but incomplete answers. Productive AI therefore requires organisations to identify relevant knowledge across both structured and unstructured sources, connect it and make it accessible under controlled conditions.

2. Valuable information can remain invisible because of its format

Not every piece of enterprise information exists in a form that can immediately be searched or analysed.

Scanned PDF files without text recognition, images containing tables, training recordings, audio files and video conferences may hold valuable operational knowledge while remaining effectively invisible to AI systems.

Technically, this barrier can already be overcome. Optical character recognition can extract information from scans and images, while transcription technologies can convert audio and video into searchable text.

The challenge is therefore less about whether such content can be processed and more about whether organisations systematically include it in their data preparation strategy. Only when previously inaccessible formats are incorporated can an AI knowledge base begin to reflect the actual information environment of the company.

3. More data does not automatically produce better AI

One of the most common assumptions in AI projects is that larger datasets will automatically improve results. In practice, the opposite can happen.

Loading all available information into a vector database or another target system can introduce outdated documents, duplicates, irrelevant files and contradictory versions. This increases the amount of noise an AI system must navigate and can reduce the reliability of its output.

There is also a direct cost dimension. Every unnecessary document requires storage, processing and potentially additional token usage.

Successful AI initiatives therefore start by asking which information is genuinely relevant to a specific use case. Data can be assessed using deterministic rules based on factors such as age, redundancy, format or sensitivity. This allows organisations to reduce their data footprint before expensive content analysis begins.

The important point is that not every document has to be processed by an AI model simply to determine whether it should be used by AI.

4. AI exposes legacy access-control problems

Generative AI can also reveal security weaknesses that have existed for years.

Many organisations deploy AI assistants without first reviewing existing permission structures. Access rights that were originally designed for small teams may have expanded over time, while shared links, inherited permissions and collaboration tools can unintentionally expose sensitive information to wider groups.

An AI assistant capable of searching across multiple repositories can make such weaknesses immediately visible. The AI itself did not create the access problem; it merely accelerates the discovery and potential exploitation of permissions that were already too broad.

Before AI systems receive production access, organisations should therefore assess ownership structures and access rights wherever the connected source provides this information. Suspicious permissions must then be reviewed and corrected by the organisation itself.

This transforms access governance from an administrative hygiene issue into a central prerequisite for secure AI deployment.

5. Sensitive information should not automatically become AI data

Creating an AI knowledge base is not purely a technical exercise. It is also a governance and compliance task.

Personal data, confidential research and development documents, financial information and other sensitive content cannot simply be made available to every AI application. Different departments also require different information domains and access levels.

The key question is therefore no longer how many systems can be connected, but which data should be used, excluded, cleaned or transformed before it reaches an AI environment.

Sensitive information can be classified using defined rules and, depending on the use case, anonymised, pseudonymised or removed before being transferred to external systems. The decisive factor is that these rules are established at the beginning of an AI project rather than added after deployment.

This is particularly important in regulated environments, where the use of AI does not remove existing obligations around confidentiality, data protection or accountability.

6. Data sovereignty is more than a hosting decision

The debate around AI sovereignty often focuses on where data is stored – in the public cloud, a private cloud, an on-premises data centre or a hybrid infrastructure.

Storage location matters, but it does not by itself create sovereignty.

An organisation is only in control of its data if it knows what information it holds, where that information resides and who owns or can access it. Without this transparency, every storage model can become a risk.

The consequences extend well beyond AI. Following a security incident or a data protection request, organisations may struggle to determine which files were affected, where copies exist and whether those files should have been processed in the first place.

Continuous discovery and scheduled rescanning of data repositories can help maintain an up-to-date picture of the information environment. Where data-inventory software runs within the organisation’s own infrastructure and the resulting index remains there as well, source data can also stay within the controlled environment during analysis.

That distinction is becoming increasingly important as organisations combine AI adoption with broader requirements around digital sovereignty and regulatory control.

Data readiness becomes an AI success factor

The broader lesson is that generative AI shifts the value of data governance.

In traditional IT environments, poorly structured data may primarily create inefficiency. In an AI environment, the same weaknesses can directly affect the reliability of automated answers, the confidentiality of information and the cost of operating the system.

The strategic question is therefore no longer simply which model an organisation should use. It is whether the organisation has prepared its information environment well enough for that model to work responsibly.

“Whether a company succeeds with AI is not decided by the choice of language model,” says Adrian Knapp, CEO of APARAVI. “What matters is the quality of the underlying data. Is the information up to date? Are permissions correct? Are sensitive contents identified? Are relevant documents even available in a format that AI systems can process?”

According to Knapp, productive AI starts long before the first prompt is entered.

“Good AI results do not begin with the model, but with a controlled and suitable data foundation. And that foundation has to exist before the first AI query is made.”

For security and technology leaders, this changes the order of priorities. The competitive advantage will not come solely from deploying the newest model fastest. It will increasingly depend on whether organisations can discover, classify, protect and continuously maintain the information on which those models rely.

In that sense, the decisive AI infrastructure of the coming years may not be the model itself.

It may be the data discipline behind it.

Reference: Challapally, Pease, Raskar & Chari (2025): The GenAI Divide – State of AI in Business 2025, MIT NANDA.

Related Articles

When Liberalism Comes Under General Suspicion

Editor’s comment: Germany’s political debate is becoming increasingly prone to a dangerous shortcut: those who argue for a smaller state, less redistribution and more market freedom can quickly find themselves placed under ideological suspicion. The controversy...

Aerial scouts are reshaping perimeter protection

Aerial scouts are reshaping perimeter protection

Drones add a capability that conventional site security has so far lacked: they can not only detect suspicious activity, but also track it across large areas. Security service provider autosecure demonstrates how unmanned aircraft can be integrated with sensors,...

Share This