---
title: "Resources — DataXray"
description: "Research, blog posts, product insights and news on classifying and governing unstructured data, newest first."
url: "https://www.dataxray.io/resources/"
language: "en-US"
---

Resources

# Research, field notes and news from DataXray

Everything DataXray publishes in one place: independent research on how organizations govern unstructured data, field notes on classifying files by what they contain rather than what they are called, product insights from real deployments, and company news. Filter by type, or read newest first.

Search resources

- [All 42](https://www.dataxray.io/resources/?)
- [Blog 19](https://www.dataxray.io/resources/?type=blog)
- [Research 2](https://www.dataxray.io/resources/?type=research)
- [News 5](https://www.dataxray.io/resources/?type=news)
- [Events & webinars 8](https://www.dataxray.io/resources/?type=events)
- [Product insights 5](https://www.dataxray.io/resources/?type=product-insights)
- [Case studies 3](https://www.dataxray.io/resources/?type=case-studies)

Showing 42

- Webinar: Harnessing Unstructured Data for Agentic AI Context

  WEBINAR

  ## [Harnessing Unstructured Data for Agentic AI Context](https://www.dataxray.io/resources/barc-webinar/)

  The files your enterprise owns are still AI's blind spot. Kevin Petrie of BARC and Kyle DuPont of DataXray on how to make them AI-ready.
- Interview: Your AI agents need context

  INTERVIEW

  ## [Your AI agents need context](https://www.dataxray.io/resources/barc-interview/)

  It's trapped in files no one can find. Kevin Petrie of BARC and Kyle DuPont of DataXray on why, and where the fix starts.
- Research: BARC research 2026

  RESEARCH

  ## [BARC research 2026](https://www.dataxray.io/research/barc-2026/)

  79% of data leaders are confident they can govern unstructured data for AI. Only 29% fully know where it lives.
- Checklist: Find it. Classify it. Get AI-ready.

  CHECKLIST

  ## [Find it. Classify it. Get AI-ready.](https://www.dataxray.io/resources/barc-checklist/)

  BARC's three-stage checklist takes your files from found to AI-ready, with the boxes to check at each stage.
- Blog: Harnessing Unstructured Data for Agentic AI: A Three-Stage Maturity Model

  BLOG

  ## [Harnessing Unstructured Data for Agentic AI: A Three-Stage Maturity Model](https://www.dataxray.io/blog/harnessing-unstructured-data-for-agentic-ai-a-three-stage-maturity-model/)

  Agentic AI runs on context, and most of an organisation's context sits in documents, emails and images it has never organised. This three-stage maturity model, baseline through established to optimised, gives data and AI teams a way to locate themselves and decide what to fix first. Written by Kevin Petrie, VP Research at BARC.
- News: 79% Confident on AI Governance. Only 29% Can Find the Data.

  NEWS

  ## [79% Confident on AI Governance. Only 29% Can Find the Data](https://www.dataxray.io/blog/barc-research-report-2026/)

  BARC surveyed 225 enterprises and found four in five confident they can use their files, emails and documents for AI without compromising governance, while fewer than one in three fully know where that data lives. The report examines what closes that gap as AI moves from experiment into production.
- Event: Ohalo at CDO Magazine Columbus Leadership Summit

  EVENT

  ## [Ohalo at CDO Magazine Columbus Leadership Summit](https://www.dataxray.io/blog/cdo-magazine-summit-foundation-problem-ai-scale/)

  At the CDO Magazine Columbus Leadership Summit, Ohalo hosted a focus group with senior data and AI leaders on why so many enterprises are still stuck in proof-of-concept. The discussion kept returning to the same cause: the files underneath. AI did not create the foundation problem, it exposed it.
- Blog: AI File Classification Accuracy: Why a 30% Miss Rate Breaks Your AI Pipeline

  BLOG

  ## [AI File Classification Accuracy: Why a 30% Miss Rate Breaks Your AI Pipeline](https://www.dataxray.io/blog/ai-file-classification-miss-rate-2026/)

  File classification accuracy measures how reliably a system identifies what is inside a file: whether it holds sensitive data, of what kind, under which policy. AI-only classifiers typically miss around 30% on complex enterprise documents, which is the number vendors rarely publish, and the reason classification belongs before ingestion rather than after.
- News: MBSD, Ohalo and Macnica Launch DSPM Implementation Service for Unstructured Data

  NEWS

  ## [MBSD, Ohalo and Macnica Launch DSPM Implementation Service for Unstructured Data](https://www.dataxray.io/blog/mbsd-ohalo-macnica-launch-dspm-service-for-unstructured-data/)

  Mitsui Bussan Secure Directions, Macnica and Ohalo launched a DSPM implementation service for unstructured data in Japan, beginning April 2026. It pairs DataXray's discovery and classification with MBSD and Macnica's deployment expertise, aimed at enterprises that must inventory, visualise and classify sensitive files before they can govern them.
- Blog: The AI Readiness Test Nobody is Passing

  BLOG

  ## [The AI Readiness Test Nobody is Passing](https://www.dataxray.io/blog/aireadiness-test-nobody-is-passing/)

  Enterprise AI is not stalling on model quality. It stalls because most of the data that would make it useful, the institutional knowledge held in emails, documents and threads, sits outside formal governance: unclassified, unauthorised, and impossible to prove safe to use. This post examines the pilot trap that results.
- Event: The AI That Wins in Financial Services Runs on Governed Unstructured Data

  EVENT

  ## [The AI That Wins in Financial Services Runs on Governed Unstructured Data](https://www.dataxray.io/blog/financial-services-unstructured-data/)

  Financial services has moved past AI experiments, yet most initiatives stall before scale. The blocker is rarely budget, intent or model maturity but unstructured data that cannot be shown to be governed. This post covers why AI readiness and compliance are the same problem, and where accountability sits between the CDO and the CIO.
- Webinar: Your Catalog Investment Works Perfectly, on 20% of Your Data.

  WEBINAR

  ## [Your Catalog Investment Works Perfectly, on 20% of Your Data](https://www.dataxray.io/blog/extend-your-data-catalog-to-unstructured-data/)

  A data catalogue works well on the structured fifth of an estate and cannot see the rest, which is where adoption quietly stalls. Drawing on deployments across financial services, manufacturing and government, this session sets out a framework for extending catalogue governance to files, and compares replacing, building and extending.
- Blog: Emerging Risks Lurk: Can Sensitive Data Classification Keep Up?

  BLOG

  ## [Emerging Risks Lurk: Can Sensitive Data Classification Keep Up?](https://www.dataxray.io/blog/emerging-risks-lurk-can-sensitive-data-classification-keep-up/)

  Most organisations classify their data and still grant far broader access than any role needs, so classification exists without changing who can reach what. This post examines where conventional classification falls short as risks shift, and what it takes for a classification to drive an access decision rather than sit in a report.
- Blog: Why DataXray Managed Metadata Sync is the Beating Heart of On-Premises DLP

  BLOG

  ## [Why DataXray Managed Metadata Sync is the Beating Heart of On-Premises DLP](https://www.dataxray.io/blog/why-ohalo-managed-metadata-sync-is-the-beating-heart-of-on-premises-dlp/)

  SharePoint Server Subscription Edition keeps data on-premises, but Microsoft's sensitivity labelling lives in cloud-only Purview, so on-premises DLP tools share no signal for which files matter. Managed Metadata Sync writes a taxonomy-driven label into SharePoint metadata itself, giving encryption, watermarking and classification tools one consistent thing to act on.
- Product insight: Streamlining Records and File Management with ROT Analysis for Enterprises

  PRODUCT INSIGHT

  ## [Streamlining Records and File Management with ROT Analysis for Enterprises](https://www.dataxray.io/blog/streamlining-records-and-file-management-with-rot-analysis-for-enterprises/)

  Redundant, obsolete and trivial files inflate storage costs and widen the attack surface, and most enterprises hold far more of them than they expect. ROT analysis finds duplicates and stale content by reading files rather than trusting their names, so records teams can delete defensibly instead of keeping everything forever.
- Blog: Reduce Data Storage Costs, Using DataXray

  BLOG

  ## [Reduce Data Storage Costs, Using DataXray](https://www.dataxray.io/blog/reduce-data-storage-costs-using-data-x-ray/)

  Storage spend grows with the estate, but a large share of what enterprises keep is redundant, obsolete or trivial and nobody can say which part. Reading files rather than trusting their names makes that measurable, so duplicates and stale content can be removed on evidence instead of guesswork.
- Case study: File-level proof for a bank divestiture

  CASE STUDY

  ## [File-level proof for a bank divestiture](https://www.dataxray.io/case-studies/bank-transforms-unstructured-data-governance-with-data-x-ray/)

  Regulators wanted file-level proof before a Tier-1 bank's divestiture could close. The bank had it in two months.
- Webinar: Transform Ungoverned Data from Compliance Risk to AI Advantage

  WEBINAR

  ## [Transform Ungoverned Data from Compliance Risk to AI Advantage](https://www.dataxray.io/blog/ai-governance-reality-making-unstructured-data-ai-ready/)

  The blocker on enterprise AI is rarely the model. It is that most of the data worth using sits in contracts, research, documents and email with no record of what is in it. This session covers how to inventory that estate, classify it, and turn ungoverned files from a compliance risk into something AI can safely use.
- Webinar: Turn SharePoint, Network Drives & More into Actionable Data

  WEBINAR

  ## [Turn SharePoint, Network Drives & More into Actionable Data](https://www.dataxray.io/blog/demo-turn-sharepoint-and-network-drives-into-actionable-data/)

  Most enterprise knowledge sits in documents nobody can query, which Gartner puts at 70-90% of all content. Extractors pull structured fields out of those files in place, across SharePoint, network drives, Box and Google Drive, turning PDFs and reports into clean JSON without moving anything. This demo follows a credit report end to end.
- Blog: Bringing Order to Data Chaos: How Ohalo is Revolutionizing Unstructured Data Security with AI

  BLOG

  ## [Bringing Order to Data Chaos: How Ohalo is Revolutionizing Unstructured Data Security with AI](https://www.dataxray.io/blog/bringing-order-to-data-chaos-how-ohalo-is-revolutionizing-unstructured-data-security-with-ai/)

  Kyle DuPont, Ohalo's CEO, joined the Cyber Security America podcast to discuss why traditional DLP struggles with unstructured data, why identity and data have become the security perimeter, and what changes once AI agents begin reading the same files. A conversation about the problem rather than a product pitch.
- Product insight: Ready or Not: AI Deep Researchers are Coming for Your Unstructured Data

  PRODUCT INSIGHT

  ## [Ready or Not: AI Deep Researchers are Coming for Your Unstructured Data](https://www.dataxray.io/blog/ready-or-not-ai-deep-researchers-are-coming-for-your-unstructured-data/)

  Deep-research AI agents read far more of an enterprise's files than a chat assistant does, and they inherit whatever access permissions already exist. Where access has drifted over years, that turns quiet oversharing into retrievable answers. This post covers what changes with agentic retrieval, and how to govern access before it arrives.
- Webinar: Live Security Session on Microsoft 365 Exposure

  WEBINAR

  ## [Live Security Session on Microsoft 365 Exposure](https://www.dataxray.io/blog/webinar-registration-june-18-oversharing-is-the-new-insider-threat/)

  SharePoint and OneDrive default to sharing rather than protecting, so files meant for a few people routinely reach a whole department. Copilot turns that latent exposure into instant retrieval: a single prompt can surface a forgotten file. This session covers how oversharing accumulates in Microsoft 365, and what to fix before an AI rollout.
- Product insight: Decoding Microsoft Purview Licensing for Automated Labeling

  PRODUCT INSIGHT

  ## [Decoding Microsoft Purview Licensing for Automated Labeling](https://www.dataxray.io/blog/decoding-microsoft-purview-licensing-for-automated-labeling/)

  Microsoft Purview's licensing for automated sensitivity labelling is genuinely hard to read: labels appear in mid-tier Office plans, but applying them automatically across SharePoint Online and OneDrive lands almost everyone on E5. This post walks the licensing matrix, the MC736438 shift toward enforcement, and what both mean for a budget.
- Product insight: Enterprise data classification in Microsoft Purview

  PRODUCT INSIGHT

  ## [Enterprise data classification in Microsoft Purview](https://www.dataxray.io/blog/enterprise-data-classification-in-microsoft-purview/)

  Microsoft Purview integrates deeply with Microsoft 365 and can scan cloud platforms such as AWS, but its limits only become obvious in a real deployment. This post records concrete results from scanning an S3 bucket and a OneDrive, the two different engines underneath, and the errors that stall multi-year rollouts.
- Blog: Unstructured Data: A Complete Guide

  BLOG

  ## [Unstructured Data: A Complete Guide](https://www.dataxray.io/blog/unstructured-data-complete-guide/)

  Unstructured data is content without a predefined model, such as documents, email, images and audio, and it accounts for roughly 90% of what enterprises hold while growing faster than structured data. This guide covers what qualifies, where it comes from, the value it carries, and the risk of leaving it unmanaged.
- Blog: Jumbled files from two merged estates flow through DataXray and emerge classified and ordered

  BLOG

  ## [Securing Unstructured Data Post-Merger: A CISO’s Guide](https://www.dataxray.io/blog/securing-unstructured-data-post-merger-a-cisos-guide/)

  A merger joins two unstructured estates with different classification standards, security controls and retention rules, and the gap between them is where risk concentrates. This guide sets out where a CISO should start: unified classification first, then regulatory obligations, then the access model a Zero Trust programme will eventually need.
- News: Ohalo Secures $3.5m Investment from YFM and Existing Investors to Propel Global Expansion

  NEWS

  ## [Ohalo Secures $3.5m Investment from YFM and Existing Investors to Propel Global Expansion](https://www.dataxray.io/blog/ohalo-secures-3-5m-investment-from-yfm-and-existing-investors-to-propel-global-expansion/)

  Ohalo raised $3.5 million in growth capital from YFM Equity Partners and existing investors, to expand internationally and develop DataXray further. Founded in 2017, the company builds tooling for mapping, classifying and redacting unstructured data so that enterprises can store, analyse and share it with less risk.
- Blog: Unveiled: The Shocking Data Chaos Cost to Your Bottom Line - Act Now

  BLOG

  ## [Unveiled: The Shocking Data Chaos Cost to Your Bottom Line - Act Now](https://www.dataxray.io/blog/infographic-the-hidden-cost-of-data-chaos/)

  Data chaos costs money in ways that rarely show up in one budget line: storage for files nobody needs, hours spent searching, compliance exposure and missed revenue from data that cannot be used. This infographic sets out where those hidden costs come from and the simple fixes that recover them.
- Blog: Decoding Supply Chain Secrets: How Unstructured Data Unleashes Efficiency

  BLOG

  ## [Decoding Supply Chain Secrets: How Unstructured Data Unleashes Efficiency](https://www.dataxray.io/blog/decoding-supply-chain-secrets-how-unstructured-data-unleashes-efficiency/)

  Unstructured data such as contracts, invoices, emails and supplier agreements holds much of a supply chain's operational insight and much of its sensitive data. Classifying those files automatically makes them faster to find, safer to share through redaction, and easier to monitor, which cuts operational errors and cost.
- Blog: Build Metadata to Streamline AI Implementation with your Own Data

  BLOG

  ## [Build Metadata to Streamline AI Implementation with your Own Data](https://www.dataxray.io/blog/build-metadata-to-streamline-ai-implementation-with-your-own-data/)

  Metadata is what makes an organisation's own documents usable by AI: without it a model cannot tell a signed contract from a draft, or restricted material from public. The second part of this series covers generating metadata automatically from file content, and what to capture for retrieval to actually work.
- Case study: 110 million files ready to move

  CASE STUDY

  ## [110 million files ready to move](https://www.dataxray.io/case-studies/enterprise-wide-data-classification-for-legacy-file-migration/)

  An energy company needed to know what was in 110 million files before moving them to a new catalog. Classification ran automatically, so the files were ready before the move.
- Blog: Auto Classification with LLMs: Turbocharging Document Classification with Generative AI

  BLOG

  ## [Auto Classification with LLMs: Turbocharging Document Classification with Generative AI](https://www.dataxray.io/blog/auto-classification-with-llms-turbocharging-document-classification-with-generative-ai/)

  Rule-based document classification breaks down on enterprise files, where the same contract arrives in a dozen shapes and no pattern covers them all. Language models read the content instead of matching patterns, which makes auto-classification workable at scale. This post covers where generative classification helps, where deterministic rules still win, and how they combine.
- Blog: The DataXray Advantage in Your DSPM Strategy

  BLOG

  ## [The DataXray Advantage in Your DSPM Strategy](https://www.dataxray.io/blog/the-data-x-ray-advantage-in-your-dspm-strategy/)

  Data Security Posture Management asks one question continuously: where is sensitive data, who can reach it, and is that still appropriate. DSPM tools answer it well for structured stores and poorly for files. This post sets out the DSPM process and where content-level classification of unstructured data fits within it.
- Blog: Mastering Unstructured Data Classification: The Key to Effective Data Governance

  BLOG

  ## [Mastering Unstructured Data Classification: The Key to Effective Data Governance](https://www.dataxray.io/blog/mastering-unstructured-data-classification-the-key-to-effective-data-governance/)

  Unstructured data classification is the practice of establishing what a file contains and how sensitive it is, so that policy can be applied to it. This post covers why classification underpins governance, the main types and methods available, and the strategies that keep a programme working as the estate keeps growing.
- Blog: What is Data Protection & Why is it Important?

  BLOG

  ## [What is Data Protection & Why is it Important?](https://www.dataxray.io/blog/what-is-data-protection-why-is-it-important/)

  Data protection is the set of controls that keep information available to the people who should have it and out of reach of everyone else. Unstructured content is the hardest part of that, making up as much as 80% of an organisation's data. This guide covers why it matters and how to approach it.
- Blog: What is Unstructured Data Discovery? Why is it Important?

  BLOG

  ## [What is Unstructured Data Discovery? Why is it Important?](https://www.dataxray.io/blog/what-is-unstructured-data-discovery-why-is-it-important/)

  Unstructured data discovery is the process of finding files across an estate and establishing what is inside them, rather than merely where they sit. It matters because the same content is both a source of insight and a regulatory exposure. This guide sets out a five-step discovery process and what to look for in a tool.
- Product insight: Enhancing Microsoft Purview Audit Trails

  PRODUCT INSIGHT

  ## [Enhancing Microsoft Purview Audit Trails](https://www.dataxray.io/blog/data-x-ray-enhances-microsoft-purview-audit-trails/)

  Microsoft Purview records file activity, but its audit trails lack the content context needed to judge whether an action mattered. Feeding DataXray's sensitivity and metadata into those trails shows not only that a file was opened but what was in it, which is the difference between an event log and an investigation.
- Blog: NLP Advances and Where it Might Take the World of Enterprise Data Security

  BLOG

  ## [NLP Advances and Where it Might Take the World of Enterprise Data Security](https://www.dataxray.io/blog/natural-language-processing-data-security/)

  Natural language processing lets software read a document rather than match patterns in it, which changes what data security can do: categorise content more accurately, inherit sensitivity between related documents, and bring a person in where the machine is unsure. This post looks at where NLP had reached, and where it was heading next.
- News: Scattered files flow through DataXray into an organized Collibra catalog

  NEWS

  ## [Ohalo, Collibra Join Forces for Unstructured Data Discovery and Cataloging](https://www.dataxray.io/blog/ohalo-collibra-join-forces/)

  Collibra catalogues and governs structured data well, but its coverage stops at the files where most enterprise content actually lives. As a Collibra Silver Partner, Ohalo feeds DataXray's file-level classifications into the catalogue, so unstructured content carries the same governance, discovery and protection as everything already inventoried.
- Case study: Safety records made safe to share

  CASE STUDY

  ## [Safety records made safe to share](https://www.dataxray.io/case-studies/improving-health-and-safety-outcomes-in-the-construction-sector/)

  Personal data kept contractors from sharing their safety records. 800,000 records were ready to share in one day.
- News: Press release: Ohalo begins work with GCHQ and NCSC

  NEWS

  ## [Press release: Ohalo begins work with GCHQ and NCSC](https://www.dataxray.io/blog/ohalo-working-with-gchq-and-ncsc/)

  Ohalo joined the GCHQ, NCSC and Wayra cybersecurity programme, which selects a small number of cybersecurity companies and pairs national security expertise with Telefonica's commercial network. The programme's aim is to make organisations easier to secure, easier to understand and defend, and quicker to detect and respond to threats.
- Blog: Using Cloud Services like Google Drive and Box and How you Should be Thinking about GDPR

  BLOG

  ## [Using Cloud Services like Google Drive and Box and How you Should be Thinking about GDPR](https://www.dataxray.io/blog/using-cloud-file-storage-and-how-to-think-about-gdpr/)

  Under GDPR, personal data in Google Drive, Box and similar cloud services has to be found, justified and secured like any other. Because these tools hold any kind of file, the only reliable way to know what personal data you hold, where it sits and who can reach it is to scan the contents of the files themselves.
