Research, field notes and news from DataXray
Everything DataXray publishes in one place: independent research on how organizations govern unstructured data, field notes on classifying files by what they contain rather than what they are called, product insights from real deployments, and company news. Filter by type, or read newest first.
Nothing matches. Try fewer words, or show all types.
WEBINAR
Harnessing Unstructured Data for Agentic AI Context
The files your enterprise owns are still AI's blind spot. Kevin Petrie of BARC and Kyle DuPont of DataXray on how to make them AI-ready.
INTERVIEW
Your AI agents need context
It's trapped in files no one can find. Kevin Petrie of BARC and Kyle DuPont of DataXray on why, and where the fix starts.
RESEARCH
BARC research 2026
79% of data leaders are confident they can govern unstructured data for AI. Only 29% fully know where it lives.
CHECKLIST
Find it. Classify it. Get AI-ready.
BARC's three-stage checklist takes your files from found to AI-ready, with the boxes to check at each stage.
BLOG
Harnessing Unstructured Data for Agentic AI: A Three-Stage Maturity Model
Agentic AI runs on context, and most of an organisation's context sits in documents, emails and images it has never organised. This three-stage maturity model, baseline through established to optimised, gives data and AI teams a way to locate themselves and decide what to fix first. Written by Kevin Petrie, VP Research at BARC.
NEWS
79% Confident on AI Governance. Only 29% Can Find the Data
BARC surveyed 225 enterprises and found four in five confident they can use their files, emails and documents for AI without compromising governance, while fewer than one in three fully know where that data lives. The report examines what closes that gap as AI moves from experiment into production.
EVENT
Ohalo at CDO Magazine Columbus Leadership Summit
At the CDO Magazine Columbus Leadership Summit, Ohalo hosted a focus group with senior data and AI leaders on why so many enterprises are still stuck in proof-of-concept. The discussion kept returning to the same cause: the files underneath. AI did not create the foundation problem, it exposed it.
BLOG
AI File Classification Accuracy: Why a 30% Miss Rate Breaks Your AI Pipeline
File classification accuracy measures how reliably a system identifies what is inside a file: whether it holds sensitive data, of what kind, under which policy. AI-only classifiers typically miss around 30% on complex enterprise documents, which is the number vendors rarely publish, and the reason classification belongs before ingestion rather than after.
NEWS
MBSD, Ohalo and Macnica Launch DSPM Implementation Service for Unstructured Data
Mitsui Bussan Secure Directions, Macnica and Ohalo launched a DSPM implementation service for unstructured data in Japan, beginning April 2026. It pairs DataXray's discovery and classification with MBSD and Macnica's deployment expertise, aimed at enterprises that must inventory, visualise and classify sensitive files before they can govern them.
BLOG
The AI Readiness Test Nobody is Passing
Enterprise AI is not stalling on model quality. It stalls because most of the data that would make it useful, the institutional knowledge held in emails, documents and threads, sits outside formal governance: unclassified, unauthorised, and impossible to prove safe to use. This post examines the pilot trap that results.
EVENT
The AI That Wins in Financial Services Runs on Governed Unstructured Data
Financial services has moved past AI experiments, yet most initiatives stall before scale. The blocker is rarely budget, intent or model maturity but unstructured data that cannot be shown to be governed. This post covers why AI readiness and compliance are the same problem, and where accountability sits between the CDO and the CIO.
WEBINAR
Your Catalog Investment Works Perfectly, on 20% of Your Data
A data catalogue works well on the structured fifth of an estate and cannot see the rest, which is where adoption quietly stalls. Drawing on deployments across financial services, manufacturing and government, this session sets out a framework for extending catalogue governance to files, and compares replacing, building and extending.
BLOG
Emerging Risks Lurk: Can Sensitive Data Classification Keep Up?
Most organisations classify their data and still grant far broader access than any role needs, so classification exists without changing who can reach what. This post examines where conventional classification falls short as risks shift, and what it takes for a classification to drive an access decision rather than sit in a report.
BLOG
Why DataXray Managed Metadata Sync is the Beating Heart of On-Premises DLP
SharePoint Server Subscription Edition keeps data on-premises, but Microsoft's sensitivity labelling lives in cloud-only Purview, so on-premises DLP tools share no signal for which files matter. Managed Metadata Sync writes a taxonomy-driven label into SharePoint metadata itself, giving encryption, watermarking and classification tools one consistent thing to act on.
PRODUCT INSIGHT
Streamlining Records and File Management with ROT Analysis for Enterprises
Redundant, obsolete and trivial files inflate storage costs and widen the attack surface, and most enterprises hold far more of them than they expect. ROT analysis finds duplicates and stale content by reading files rather than trusting their names, so records teams can delete defensibly instead of keeping everything forever.
BLOG
Reduce Data Storage Costs, Using DataXray
Storage spend grows with the estate, but a large share of what enterprises keep is redundant, obsolete or trivial and nobody can say which part. Reading files rather than trusting their names makes that measurable, so duplicates and stale content can be removed on evidence instead of guesswork.
CASE STUDY
File-level proof for a bank divestiture
Regulators wanted file-level proof before a Tier-1 bank's divestiture could close. The bank had it in two months.
WEBINAR
Transform Ungoverned Data from Compliance Risk to AI Advantage
The blocker on enterprise AI is rarely the model. It is that most of the data worth using sits in contracts, research, documents and email with no record of what is in it. This session covers how to inventory that estate, classify it, and turn ungoverned files from a compliance risk into something AI can safely use.
WEBINAR
Turn SharePoint, Network Drives & More into Actionable Data
Most enterprise knowledge sits in documents nobody can query, which Gartner puts at 70-90% of all content. Extractors pull structured fields out of those files in place, across SharePoint, network drives, Box and Google Drive, turning PDFs and reports into clean JSON without moving anything. This demo follows a credit report end to end.
BLOG
Bringing Order to Data Chaos: How Ohalo is Revolutionizing Unstructured Data Security with AI
Kyle DuPont, Ohalo's CEO, joined the Cyber Security America podcast to discuss why traditional DLP struggles with unstructured data, why identity and data have become the security perimeter, and what changes once AI agents begin reading the same files. A conversation about the problem rather than a product pitch.
PRODUCT INSIGHT
Ready or Not: AI Deep Researchers are Coming for Your Unstructured Data
Deep-research AI agents read far more of an enterprise's files than a chat assistant does, and they inherit whatever access permissions already exist. Where access has drifted over years, that turns quiet oversharing into retrievable answers. This post covers what changes with agentic retrieval, and how to govern access before it arrives.
WEBINAR
Live Security Session on Microsoft 365 Exposure
SharePoint and OneDrive default to sharing rather than protecting, so files meant for a few people routinely reach a whole department. Copilot turns that latent exposure into instant retrieval: a single prompt can surface a forgotten file. This session covers how oversharing accumulates in Microsoft 365, and what to fix before an AI rollout.
PRODUCT INSIGHT
Decoding Microsoft Purview Licensing for Automated Labeling
Microsoft Purview's licensing for automated sensitivity labelling is genuinely hard to read: labels appear in mid-tier Office plans, but applying them automatically across SharePoint Online and OneDrive lands almost everyone on E5. This post walks the licensing matrix, the MC736438 shift toward enforcement, and what both mean for a budget.
PRODUCT INSIGHT
Enterprise data classification in Microsoft Purview
Microsoft Purview integrates deeply with Microsoft 365 and can scan cloud platforms such as AWS, but its limits only become obvious in a real deployment. This post records concrete results from scanning an S3 bucket and a OneDrive, the two different engines underneath, and the errors that stall multi-year rollouts.
BLOG
Unstructured Data: A Complete Guide
Unstructured data is content without a predefined model, such as documents, email, images and audio, and it accounts for roughly 90% of what enterprises hold while growing faster than structured data. This guide covers what qualifies, where it comes from, the value it carries, and the risk of leaving it unmanaged.
BLOG
Securing Unstructured Data Post-Merger: A CISO’s Guide
A merger joins two unstructured estates with different classification standards, security controls and retention rules, and the gap between them is where risk concentrates. This guide sets out where a CISO should start: unified classification first, then regulatory obligations, then the access model a Zero Trust programme will eventually need.
NEWS
Ohalo Secures $3.5m Investment from YFM and Existing Investors to Propel Global Expansion
Ohalo raised $3.5 million in growth capital from YFM Equity Partners and existing investors, to expand internationally and develop DataXray further. Founded in 2017, the company builds tooling for mapping, classifying and redacting unstructured data so that enterprises can store, analyse and share it with less risk.
BLOG
Unveiled: The Shocking Data Chaos Cost to Your Bottom Line - Act Now
Data chaos costs money in ways that rarely show up in one budget line: storage for files nobody needs, hours spent searching, compliance exposure and missed revenue from data that cannot be used. This infographic sets out where those hidden costs come from and the simple fixes that recover them.
BLOG
Decoding Supply Chain Secrets: How Unstructured Data Unleashes Efficiency
Unstructured data such as contracts, invoices, emails and supplier agreements holds much of a supply chain's operational insight and much of its sensitive data. Classifying those files automatically makes them faster to find, safer to share through redaction, and easier to monitor, which cuts operational errors and cost.
BLOG
Build Metadata to Streamline AI Implementation with your Own Data
Metadata is what makes an organisation's own documents usable by AI: without it a model cannot tell a signed contract from a draft, or restricted material from public. The second part of this series covers generating metadata automatically from file content, and what to capture for retrieval to actually work.
CASE STUDY
110 million files ready to move
An energy company needed to know what was in 110 million files before moving them to a new catalog. Classification ran automatically, so the files were ready before the move.
BLOG
Auto Classification with LLMs: Turbocharging Document Classification with Generative AI
Rule-based document classification breaks down on enterprise files, where the same contract arrives in a dozen shapes and no pattern covers them all. Language models read the content instead of matching patterns, which makes auto-classification workable at scale. This post covers where generative classification helps, where deterministic rules still win, and how they combine.
BLOG
The DataXray Advantage in Your DSPM Strategy
Data Security Posture Management asks one question continuously: where is sensitive data, who can reach it, and is that still appropriate. DSPM tools answer it well for structured stores and poorly for files. This post sets out the DSPM process and where content-level classification of unstructured data fits within it.
BLOG
Mastering Unstructured Data Classification: The Key to Effective Data Governance
Unstructured data classification is the practice of establishing what a file contains and how sensitive it is, so that policy can be applied to it. This post covers why classification underpins governance, the main types and methods available, and the strategies that keep a programme working as the estate keeps growing.
BLOG
What is Data Protection & Why is it Important?
Data protection is the set of controls that keep information available to the people who should have it and out of reach of everyone else. Unstructured content is the hardest part of that, making up as much as 80% of an organisation's data. This guide covers why it matters and how to approach it.
BLOG
What is Unstructured Data Discovery? Why is it Important?
Unstructured data discovery is the process of finding files across an estate and establishing what is inside them, rather than merely where they sit. It matters because the same content is both a source of insight and a regulatory exposure. This guide sets out a five-step discovery process and what to look for in a tool.
PRODUCT INSIGHT
Enhancing Microsoft Purview Audit Trails
Microsoft Purview records file activity, but its audit trails lack the content context needed to judge whether an action mattered. Feeding DataXray's sensitivity and metadata into those trails shows not only that a file was opened but what was in it, which is the difference between an event log and an investigation.
BLOG
NLP Advances and Where it Might Take the World of Enterprise Data Security
Natural language processing lets software read a document rather than match patterns in it, which changes what data security can do: categorise content more accurately, inherit sensitivity between related documents, and bring a person in where the machine is unsure. This post looks at where NLP had reached, and where it was heading next.
NEWS
Ohalo, Collibra Join Forces for Unstructured Data Discovery and Cataloging
Collibra catalogues and governs structured data well, but its coverage stops at the files where most enterprise content actually lives. As a Collibra Silver Partner, Ohalo feeds DataXray's file-level classifications into the catalogue, so unstructured content carries the same governance, discovery and protection as everything already inventoried.
CASE STUDY
Safety records made safe to share
Personal data kept contractors from sharing their safety records. 800,000 records were ready to share in one day.
NEWS
Press release: Ohalo begins work with GCHQ and NCSC
Ohalo joined the GCHQ, NCSC and Wayra cybersecurity programme, which selects a small number of cybersecurity companies and pairs national security expertise with Telefonica's commercial network. The programme's aim is to make organisations easier to secure, easier to understand and defend, and quicker to detect and respond to threats.
BLOG
Using Cloud Services like Google Drive and Box and How you Should be Thinking about GDPR
Under GDPR, personal data in Google Drive, Box and similar cloud services has to be found, justified and secured like any other. Because these tools hold any kind of file, the only reliable way to know what personal data you hold, where it sits and who can reach it is to scan the contents of the files themselves.