Skip to content
Ohalo is now DataXray.What changed
Platform

Read, classify and act on every file

DataXray is one platform for every file in your estate. Native connectors reach where your files live, layered classification reads what is inside each one, and a single searchable index feeds the security, governance, and AI tools you already run.

How the platform fits together

Deploy DataXray where your data lives

DataXray installs inside your environment, including closed, air-gapped enclaves. On its own, it takes every file from discovery to classification and feeds the results into your stack. Add-on modules for records, disclosure, legal review, AI access, and taxonomy run on top of it.

Your environment

  • On-premises
  • Cloud
  • Hybrid
  • Air-gapped
  • Agentless
  • Containerized
  • Operational in hours
  • Your data never moves
  • Language models run in-boundary

The platform

DataXray
  1. Connect to every source
  2. Read every file
  3. Classify what's inside
  4. Feed your stack

Modules

How DataXray works

DataXray works in four stages, and each one feeds the next.

  1. 01DiscoverDataXray connects to the systems where your files live and builds an inventory of every file and its technical metadata across your estate. Discovery is agentless and read-only.
  2. 02ReadIt opens each file and extracts the text, including from scanned images with OCR, nested archives, and email attachments. It handles hundreds of file types in over 100 languages.
  3. 03ClassifyIt works out what each file contains, such as personal data, payment card numbers, health records, controlled unclassified information, or intellectual property. Fixed rules, machine-learning models, and language models work together to read each document in context, not just match patterns. You can add your own categories, and human-in-the-loop review checks the results.
  4. 04ActYou can turn the results into retention labels or redaction, and DataXray feeds file-level attributes into the security and governance tools you already run, which handle encryption and access.

What comparable tools do, and what DataXray does instead

Data-mapping and DLP tools work out where sensitive data probably lives from its location and metadata, and stop at a map or an alert. DataXray opens the file, works out what is actually in it, and acts on that. A label, a redaction, or a retention rule can only be applied with confidence when it rests on what the file contains.

Evidence

1BN+ files under ongoing governance

98.7%
Classification success rate
Measured on document-level PII/PCI detection in text-based English files.
72,000
Files per hour
DataXray reads content at 100,000s of words per second.
55+
Native connectors
Plus a universal connector where no native connector exists.

Capabilities

When DataXray scans your estate, it keeps a record of every file, its metadata, and what it found inside, all in one place you can search. Every capability below works from that record, so a file only needs to be scanned and classified once to serve your security, compliance, governance, and AI teams.

  • One searchable indexEvery file, its metadata, its annotations and its classifications land in one index you can search and filter by sensitivity, person or category.
  • Classifies with layered intelligencePredefined annotator packs cover regulations such as GDPR and CPRA, and you can build custom classifiers and extractors for your own categories.
  • Extracts structured data from documentsThe LLM extractor turns file contents into structured, queryable fields for analytics and retrieval pipelines.
  • Redacts sensitive contentRedact, anonymize, or pseudonymize text for disclosure, sharing, or release.
  • Applies retention and detects ROTRetention labels, and detection of redundant, obsolete and trivial data.
  • Pushes classifications into your stackPushes file-level attributes into Microsoft Purview MIP, Virtru, Netskope, Box Shield, Collibra, Thales, and Atlan. With Virtru, protection travels with the file.
  • Governs what AI may seePre-ingestion classification for RAG and agent pipelines, so models index only what they should.
  • Supports DSARs, legal holds and regulatory evidenceFiles can be gathered into case files, archived for legal hold, or exported as reports. Every finding is logged with what was found, which rule or model found it, and when the file was last checked, so the audit trail holds up.
DataXray integrations

Connects to where files live, feeds what protects them

DataXray enriches the stack you already have rather than replacing it. It ships with 55+ native connectors to datasource types plus a universal connector: where no native connector exists, a Python SDK lets your agents build a RPC connection to the source in days. A REST API supports your own downstream processes. See Integrations.

Deployment

Runs where the data is. DataXray is agentless and containerized. It deploys to cloud, hybrid, on-premises or fully air-gapped environments on RHEL with Podman or Docker, and is operational in hours.

Your data never leaves your environment. DataXray's language models run on your own infrastructure, so your files are never sent out for analysis. DataXray connects to your sources with read-only access by default, using accounts limited to the access DataXray needs.

Enterprise and government

The certifications your vendor security review will check

  • DoW ATOAuthority to Operate on three networks in the U.S. Department of War
  • SOC 2 Type 2Ohalo Limited, annual audit
  • ISO 27001Information security management
  • FIPSFIPS-compliant cryptography
  • STIGSTIG-compliant implementations

Frequently asked questions

What does DataXray do?

DataXray is the unstructured intelligence layer for the enterprise. It finds every file across your estate, reads what is inside, and classifies sensitive or business-critical data at the file level. You can then apply retention labels or redaction, and the results flow into the security, governance, and AI tools you already run.

Do I need a module to get value from DataXray?

No. DataXray works on its own, from discovery to action. Modules are add-ons for teams that need a dedicated workflow, such as governed AI assistant access, records management, or litigation review. They work on the files DataXray has already scanned and classified, so adding one needs no new scan.

How accurate is DataXray's classification?

DataXray has a 98.7% classification success rate, measured on document-level PII and PCI detection in text-based English files. Fixed rules, machine-learning models, and language models read each document in context, working with your own rules and human-in-the-loop review.

How fast does it scan?

DataXray reads hundreds of thousands of words per second and processes around 1,200 files per minute, or 72,000 files per hour. The pipeline scales horizontally, so throughput can be increased for petabyte-scale estates.

Does our data leave our environment?

No. DataXray deploys inside your environment, including closed, air-gapped enclaves. Your data never leaves it, and DataXray's language models run on your own infrastructure.

What if you do not have a connector for our system?

DataXray has 55+ native connectors. Where no native connector exists, the universal connector lets a custom connector be built, so legacy, bespoke, and in-house systems are in scope rather than an exception.

In 30 minutes, see how DataXray reads what other tools only label