The same software, wherever your data has to stay
DataXray deploys inside your environment: on-premises, in your own cloud, hybrid, or in closed, air-gapped enclaves with no external connectivity. It is agentless and containerized, runs on RHEL with Podman or Docker, and is operational in hours. Language models run in-boundary, sources are read-only, and customer data never leaves your environment.
Deployment options
DataXray runs as the same software in every environment, so there is no reduced-capability edition for disconnected networks and no feature that depends on reaching a hosted service. What changes between deployments is where the containers run and how images arrive, not what the platform can do.
- On-premisesPhysical or virtual servers in your own data center, reading your file servers, NAS and on-premises collaboration platforms.
- Private cloudDeployed in your own cloud tenancy, reading cloud file and object storage, Microsoft 365 and Google Workspace alongside anything on-premises.
- HybridOne deployment spanning on-premises and cloud sources, with results in one index.
- Air-gappedClosed enclaves with zero external connectivity. Container images are pulled once and the deployment runs with no external connection thereafter.
Architecture
DataXray sits inside your boundary between your data sources and your enforcement tools. Connectors read files read-only, the platform extracts and classifies their contents with in-boundary models, the results land in an index inside your environment, and classifications flow out to the labeling, encryption, DLP and catalog tools you already run.
- Connectors55+ native connectors to datasource types, plus a universal connector: where no native connector exists, AI builds one for your source in days. Native connectors recrawl intelligently, so only changed content is re-read. See Integrations.
- Extraction and classificationFiles are opened and their text extracted, including scanned images through OCR, nested archives and email attachments, then classified by layered word lists, patterns, NLP, machine-learning annotators, LLM categorization and your own rules.
- IndexEvery file, its metadata and its classifications are written to one searchable index inside your boundary. Nothing is transmitted to Ohalo.
- Actions and integrationsLabels, retention, redaction and file-level attributes pushed to Microsoft Purview MIP, Virtru, Netskope, Box Shield, Collibra, Thales and Atlan. A REST API and Python SDK support your own downstream processes.
- ModulesAI Platform, Compass, Curator, AutoFOIA and eDiscovery run on the same index, permission model and audit trail.
Security architecture
The controls a security architect will ask about first are where data goes, what the platform can change, who can sign in and what evidence it leaves. DataXray keeps data in your boundary, never mutates sources during discovery, authenticates through your identity provider and writes an audit trail your SIEM can ingest.
- AgentlessNothing is installed on endpoints or file servers.
- Read-only, least privilegeSources are accessed read-only with least-privilege service accounts. Actions such as labeling, redaction or encryption are explicit, configured steps that produce a record.
- In-boundary modelsClassification and language models run inside your environment. No inference call leaves the boundary.
- Your identity providerUsers authenticate through your identity provider, so access follows the roles you already manage.
- Audit logsEach scan produces audit logs your SIEM can ingest, and every file carries what was found, which rule or model found it and when it was last re-checked.
- Diagnostics without contentWhere support diagnostics are shared, they cover product performance, configuration and error information, not the contents of your files.
- Hardening and cryptographySTIG-compliant implementations and FIPS-compliant cryptography. See Security.
Air-gapped and classified enclaves
Most classification platforms are cloud-native and cannot scan without reaching a hosted service. DataXray was built for the opposite case: it runs in closed enclaves with zero external connectivity, with Authority to Operate achieved on three networks in the U.S. Department of War, including air-gapped ones.
- No outbound dependencyContainer images are pulled once; thereafter the deployment, its models and its index run with no external connection.
- Module containers built for disconnectionAutoFOIA ships its interface, API, database and document conversion in a single self-contained container, with a FIPS-capable runtime, encryption at rest, a classification banner, idle session lock and system-use notification. Curator runs as a single container with no telemetry, one instance per network.
- AccreditationAn ATO is granted by the customer agency for a specific system in a specific environment. STIG-compliant deployments are typically what an accreditation package requires.
Scale and performance
DataXray reads content at hundreds of thousands of words per second, around 1,200 files per minute or 72,000 files per hour, and the pipeline is horizontally scalable, so throughput can be increased for petabyte-scale estates. Curator is sized for estates of hundreds of millions of files.
Time to value
DataXray is operational in hours rather than weeks, including in air-gapped environments: pull the containers, configure sources, scan. Because it is agentless, there is no endpoint rollout, and modules such as AutoFOIA deploy in days rather than as a multi-month implementation project.
Frequently asked questions
What does DataXray run on?
RHEL with Podman or Docker, on physical or virtual infrastructure.
Does DataXray need internet access?
No. Container images are pulled once, and the deployment runs with zero external connectivity thereafter. Language models run in-boundary, so no inference call leaves your environment.
Does DataXray install agents?
No. DataXray is agentless. It reads sources over their own interfaces, read-only, with least-privilege service accounts.
Does any data go to Ohalo?
No. Files are read where they sit and the index is written inside your boundary. Support diagnostics, where shared, cover product performance, configuration and errors, not file contents.
How long does deployment take?
Hours rather than weeks, including in air-gapped environments.