Deploy on infrastructure you control.
On premises, in your cloud or air-gapped.
DataWalk runs graph, relational and search analysis in one engine over a shared data state. Deploy it within your infrastructure and access controls, with application, integration, and compute services that scale independently.
Why bring data into DataWalk?
Repeated relationship analysis should not require rebuilding the same connections for every query.
DataWalk ingests selected data and stores its relationships for reuse. Analysts can follow those relationships, filter the results, and aggregate across them in the same engine.
That preparation requires storage, ingestion, and refresh work. Your source systems stay in place. The choice is whether preparing relationships once is useful for the analysis you need to repeat.
Meaning and relationships are stored with the data.
Ontology-first. Map sources to business meaning.
Semantic ETL maps incoming records to defined entities, attributes, and relationships. New sources extend the business data model instead of leaving each application to interpret their fields.
Enterprise Object Graph. Relationships are part of the model.
DataWalk’s Enterprise Object Graph (EOG) stores typed objects and relationships prepared for traversal. Queries reuse those connections instead of reconstructing them from source-table joins each time.
Multi-model. One engine, one analytical state.
Graph, relational and search operations use the same objects and identifiers. A search result can become the starting population for a traversal, then a relational aggregation, without transfers between analytical engines.
Set traversal. Each step operates on a population.
The engine works over normalized sets. Each filter or traversal produces the population for the next step, so an analysis can narrow, expand, or change direction from its current result.
Graph consistency. Refresh follows dependencies.
Rules define derived relationships and calculations. DataWalk tracks their dependencies and refreshes affected results in order when inputs change. Permissions are enforced during queries.
Fewer systems to update when a definition changes.
Change what “active customer” means. The work depends on how many places store that definition.
Three representations. Changes across the stack.
Where each engine holds its own representation, the change must reach each one. Definitions embedded in schemas, views, and query logic can also require updates to pipelines, scores, and applications. Then the results need to be reconciled.
One shared definition. Dependencies tracked in-platform.
Update the definition in the shared business model. DataWalk tracks and refreshes dependent calculations. All three analytical methods use the updated state, without a separate propagation pipeline between engines.
Source mappings and downstream applications may still need changes. The reduction is in maintaining and reconciling separate analytical representations.
Data model and deployment details
A relationship can carry its own context. Employment can connect a person and an organization, hold a role and employment dates, and connect to a manager. Those facts remain available when the analysis follows that relationship.
The business data model defines these concepts and relationships. It is separate from machine learning models that generate predictions. Source ingestion and dependency refresh still determine data freshness.
Application services handle UI and API requests. Integration services run apps and external integrations. Database services execute analytical work in a distributed columnar engine. Each tier can scale separately.
A shared analytical state does not mean one server or one physical copy. Local caches, backups, source records, and exports still exist. Distributed matching also moves data between nodes.
Storage grows independently of query memory.
The entire graph does not need to fit in RAM.
Object storage holds persistent data. Workers cache data on local disk and use memory for queries, imports, and refresh jobs. Size storage for the data and compute resources for the workload.
This separates stored graph size from execution memory. Concurrent queries and background jobs still require sufficient CPU and RAM.
Object storage
Persistent data
Worker-local disk
Cached data
CPU and RAM
Queries, imports, and refresh jobs
Storage capacity and execution capacity are separate sizing decisions.
objects
Reported in DataWalk’s architecture presentation. This is a loaded-data scale figure, not a traversal latency or concurrent-throughput result.
Bounded traversal stayed subsecond as depth increased.
(placeholder - The September 23, 2026) benchmark found essentially flat latency across depths 2 through 15 when each step retained the 100 highest-ranked accounts. Across the four cases on the 20 TB graph, median request latency over those depths was 0.18 to 0.41 seconds.
Single user, warmed local storage, fixed six-node cluster. Each timing covers the complete request to the specified depth, not an incremental hop. The frontier limit excludes qualifying accounts from further traversal. Unbounded results are reported separately.
Traversal conditions, unbounded results, and entity resolution scaling
Directed traversal at depth 15
On the 20 TB graph, a directed, date-filtered request reached 45.6 million accounts at depth 15 in 5.34 seconds. This measures the complete traversal and count. It excludes retrieval of the account records.
At depth 7, bidirectional requests took 0.44 seconds with a date filter and 61.60 seconds without it. They reached different populations.
Entity resolution on 3 to 12 nodes
The architecture deck reports 162 million organizations and 3.7 billion candidate pairs processed in 7 hours 3 minutes on 3 nodes and 3 hours 4 minutes on 12 nodes.
Four times the nodes produced a 2.30× speedup, or about 58% scaling efficiency. Matching pairs across nodes requires a data reshuffle.
DataWalk ran the September 23, 2026 traversal benchmark on synthetic banking graphs of 1.5, 10, and 20 TB. Each contained 100 million accounts, including 50 million traversable internal accounts, with up to 20 billion internal bookings. Results are medians of 10 runs with one virtual user and warmed local storage.
Each cluster had 6 nodes, 288 vCPU, and approximately 2.25 TiB of RAM. Queries stayed within a 12 GB query-memory budget. Timings exclude record retrieval and application processing.
Account count and node count were fixed. The traversal test did not measure concurrent throughput, growth in account count, or performance as nodes are added. The entity resolution results apply to a separate workload.
The bounded traversal tests retained the 100 highest-ranked accounts after each step. Accounts outside that set were excluded from further traversal.
Sizing formulas and a worked capacity example
Object storage
The 5.1+ sizing guide estimates live database storage at 3× structured input, plus two full backups. Add capacity for unstructured content and its two backups separately.
Local disk per worker
Estimated disk per worker = 1.4 × 10 × initial data load ÷ compute servers. This includes cache expansion and a 40% reserve for growth and uneven distribution.
CPU and RAM
Size for query scope, concurrent users, imports, and overlapping refresh jobs. Apply the minimum requirements for the selected release.
Example: 250 GB input, 3 production workers
| Live database object storage | 750 GB |
| Two full database backups | 1.5 TB |
| Calculated local disk per worker | ≈1.17 TB |
| RAM per worker | 256 GB |
| Total RAM across 3 workers | 768 GB |
| Total CPU across 3 workers | 96 vCPU |
Object storage: 2.25 TB before content and growth. Local disk: 1.4 × 10 × 250 ÷ 3 ≈ 1,167 GB per worker; apply the release minimum if higher.
RAM and CPU use the documented production baseline of 256 GB and 32 vCPU per worker. They are not calculated from the 250 GB input volume. Query scope, concurrency, imports, and refresh jobs determine whether more capacity is needed.
Run it within your infrastructure and operating model.
Your cloud
Deploy in your cloud environment with supported object storage. Configure networking, identity, and access within your infrastructure.
On premises
Run on your Kubernetes infrastructure with a supported object store and worker-local disk.
Air-gapped
Use an internal image registry and controlled transfer of software, dependencies, and updates.
Your team operates the platform.
Customer administrators manage infrastructure, access, monitoring, backups, restores, and upgrades. DataWalk provides deployment guidance and administrator training as part of the handoff.
Installation establishes the environment. Ongoing operations keep it available and recoverable. Use-case implementation covers source mapping, business definitions, and analytical workflows. Assign owners for each.
Installation
Establishes the environment.
Ongoing operations
Keep it available and recoverable.
Use-case implementation
Covers source mapping, business definitions, and analytical workflows.
Environment prerequisites and compatibility
Prerequisites
- An operational Kubernetes cluster and Helm administration skills.
- At least 3 workers. The documented production baseline is 32 vCPU and 256 GB RAM per worker.
- Supported object storage and local disks meeting capacity, throughput, and IOPS requirements.
- Configured identity integration, certificates, networking, and access to deployment packages.
Confirm supported versions and disk requirements for the selected release. The minimum-requirements guide places workers in one Availability Zone. Define disaster recovery separately.
Development, test, and production
Each production instance requires at least one nonproduction environment; two are recommended. Test should match production. Development can be smaller.
Development, test, and production can use separate namespaces and Helm releases on one cluster. Allocate capacity, storage paths, ports, and monitoring for each.
The DevOps team must attend a deployment workshop. CKA or OpenShift administrator expertise is strongly recommended.
The supplied object-storage compatibility list includes Dell ECS, NetApp StorageGRID, Pure FlashBlade, and MinIO. Confirm support for the selected release.
Administrator handoff, monitoring, and upgrades
Training and installation
The handoff includes Deployment Certification and administrator training. Customer administrators must demonstrate installation and maintenance using DataWalk images and Helm charts.
Monitoring and recovery
Grafana and Prometheus provide monitoring. Configure alerts and backups to object storage. Test restores on separate infrastructure and disaster recovery against recovery point and recovery time objectives (RPO/RTO).
Upgrades
The documented procedure requires a database backup and Vault snapshot before replacing the Helm release with the Revive policy. Persistent data stays in object storage. Verify the service and recovery procedure after the upgrade.
Record the owners of infrastructure, backups, restores, and upgrades in the handoff. Define responsibilities for use-case implementation separately.
Review the evidence and the workloads.
Traversal benchmark
Test conditions, bounded and unbounded workloads, and complete request timings.
Entity resolution
How matching produces resolved entities for use across the platform.
Connected analytics
The graph, relational and search operations the architecture supports.
Questions for the architecture review.
Is DataWalk a federated query layer?
No. Data for analysis is ingested and prepared in DataWalk. Existing warehouses and operational systems can remain in place as sources.
Does a shared state mean one server or one physical copy?
No. Compute runs across workers. The platform also uses local caches and backups, and sources and downstream exports remain separate.
How quickly do source changes reach an analysis?
After ingestion and the required dependent calculations finish. Set freshness targets for that full process, including downstream refreshes.
Does the 12 GB query budget describe a production deployment?
No. It is a query-memory budget from the traversal benchmark. Each test cluster had approximately 2.25 TiB of provisioned RAM. Production capacity depends on the workload and release requirements.
Review the architecture against your workload.
Bring representative queries, data volumes, freshness targets, and deployment constraints. Review capacity, operating responsibilities, and recovery requirements.
Architecture review checklist
Review inputs and outputs
- Inputs
- Initial load, growth forecast, content volume, representative queries, concurrency, refresh targets, and deployment constraints.
- Tests
- Query completeness and latency, cold and warm storage, ingestion overlap, permission changes, restores, and disaster recovery.
- Decisions
- Topology, versions, capacity, environments, RPO/RTO targets, and operating owners.
- Outputs
- A sizing and deployment plan, measured workload limits, and an administrator handoff checklist.