Skip to content
Platform / Architecture & Deployment

Deploy on infrastructure you control.

On premises, in your cloud or air-gapped.

DataWalk runs graph, relational and search analysis in one engine over a shared data state. Deploy it within your infrastructure and access controls, with application, integration, and compute services that scale independently.

YOUR INFRASTRUCTURE · KUBERNETES Application servicesUI and APIs Integration servicesApps and source systems Distributed computeGraph · Relational · Search Shared data, business definitions, and permissionsEntity resolution and dependent calculations run in the same database. Local disk cacheData used by workers Object storagePersistent data
The challenge

Why bring data into DataWalk?

Repeated relationship analysis should not require rebuilding the same connections for every query.

DataWalk ingests selected data and stores its relationships for reuse. Analysts can follow those relationships, filter the results, and aggregate across them in the same engine.

That preparation requires storage, ingestion, and refresh work. Your source systems stay in place. The choice is whether preparing relationships once is useful for the analysis you need to repeat.

PREPARE RELATIONSHIPS FOR REPEATED ANALYSIS Selected source data INGEST AND REFRESH Stored objects and relationships Shared business definitions Graph Relational Search
The platform

Meaning and relationships are stored with the data.

01

Ontology-first. Map sources to business meaning.

Semantic ETL maps incoming records to defined entities, attributes, and relationships. New sources extend the business data model instead of leaving each application to interpret their fields.

02

Enterprise Object Graph. Relationships are part of the model.

DataWalk’s Enterprise Object Graph (EOG) stores typed objects and relationships prepared for traversal. Queries reuse those connections instead of reconstructing them from source-table joins each time.

03

Multi-model. One engine, one analytical state.

Graph, relational and search operations use the same objects and identifiers. A search result can become the starting population for a traversal, then a relational aggregation, without transfers between analytical engines.

04

Set traversal. Each step operates on a population.

The engine works over normalized sets. Each filter or traversal produces the population for the next step, so an analysis can narrow, expand, or change direction from its current result.

05

Graph consistency. Refresh follows dependencies.

Rules define derived relationships and calculations. DataWalk tracks their dependencies and refreshes affected results in order when inputs change. Permissions are enforced during queries.

Fewer systems to update when a definition changes.

Change what “active customer” means. The work depends on how many places store that definition.

Separate analytical engines

Three representations. Changes across the stack.

GraphRelationalSearch

Where each engine holds its own representation, the change must reach each one. Definitions embedded in schemas, views, and query logic can also require updates to pipelines, scores, and applications. Then the results need to be reconciled.

DataWalk

One shared definition. Dependencies tracked in-platform.

Graph, relational and search

Update the definition in the shared business model. DataWalk tracks and refreshes dependent calculations. All three analytical methods use the updated state, without a separate propagation pipeline between engines.

Source mappings and downstream applications may still need changes. The reduction is in maintaining and reconciling separate analytical representations.

Data model and deployment details

A relationship can carry its own context. Employment can connect a person and an organization, hold a role and employment dates, and connect to a manager. Those facts remain available when the analysis follows that relationship.

The business data model defines these concepts and relationships. It is separate from machine learning models that generate predictions. Source ingestion and dependency refresh still determine data freshness.

Application services handle UI and API requests. Integration services run apps and external integrations. Database services execute analytical work in a distributed columnar engine. Each tier can scale separately.

A shared analytical state does not mean one server or one physical copy. Local caches, backups, source records, and exports still exist. Distributed matching also moves data between nodes.

In detail

Storage grows independently of query memory.

The entire graph does not need to fit in RAM.

Object storage holds persistent data. Workers cache data on local disk and use memory for queries, imports, and refresh jobs. Size storage for the data and compute resources for the workload.

This separates stored graph size from execution memory. Concurrent queries and background jobs still require sufficient CPU and RAM.

Separate sizing decisions

Object storage

Persistent data

Worker-local disk

Cached data

CPU and RAM

Queries, imports, and refresh jobs

Storage capacity and execution capacity are separate sizing decisions.

Loaded and indexed in one environment 495 billion

objects

Reported in DataWalk’s architecture presentation. This is a loaded-data scale figure, not a traversal latency or concurrent-throughput result.

Median request latency · depths 2 to 15 · 20 TB graph 0.18-0.41 s

Bounded traversal stayed subsecond as depth increased.

(placeholder - The September 23, 2026) benchmark found essentially flat latency across depths 2 through 15 when each step retained the 100 highest-ranked accounts. Across the four cases on the 20 TB graph, median request latency over those depths was 0.18 to 0.41 seconds.

Single user, warmed local storage, fixed six-node cluster. Each timing covers the complete request to the specified depth, not an incremental hop. The frontier limit excludes qualifying accounts from further traversal. Unbounded results are reported separately.

Traversal conditions, unbounded results, and entity resolution scaling
Traversal test · fixed six-node clusters 5.34 s

Directed traversal at depth 15

On the 20 TB graph, a directed, date-filtered request reached 45.6 million accounts at depth 15 in 5.34 seconds. This measures the complete traversal and count. It excludes retrieval of the account records.

At depth 7, bidirectional requests took 0.44 seconds with a date filter and 61.60 seconds without it. They reached different populations.

Entity resolution test · fixed workload 2.30× faster

Entity resolution on 3 to 12 nodes

The architecture deck reports 162 million organizations and 3.7 billion candidate pairs processed in 7 hours 3 minutes on 3 nodes and 3 hours 4 minutes on 12 nodes.

Four times the nodes produced a 2.30× speedup, or about 58% scaling efficiency. Matching pairs across nodes requires a data reshuffle.

DataWalk ran the September 23, 2026 traversal benchmark on synthetic banking graphs of 1.5, 10, and 20 TB. Each contained 100 million accounts, including 50 million traversable internal accounts, with up to 20 billion internal bookings. Results are medians of 10 runs with one virtual user and warmed local storage.

Each cluster had 6 nodes, 288 vCPU, and approximately 2.25 TiB of RAM. Queries stayed within a 12 GB query-memory budget. Timings exclude record retrieval and application processing.

Account count and node count were fixed. The traversal test did not measure concurrent throughput, growth in account count, or performance as nodes are added. The entity resolution results apply to a separate workload.

The bounded traversal tests retained the 100 highest-ranked accounts after each step. Accounts outside that set were excluded from further traversal.

Sizing formulas and a worked capacity example

Object storage

The 5.1+ sizing guide estimates live database storage at 3× structured input, plus two full backups. Add capacity for unstructured content and its two backups separately.

Local disk per worker

Estimated disk per worker = 1.4 × 10 × initial data load ÷ compute servers. This includes cache expansion and a 40% reserve for growth and uneven distribution.

CPU and RAM

Size for query scope, concurrent users, imports, and overlapping refresh jobs. Apply the minimum requirements for the selected release.

Storage estimate · 5.1+ sizing guide

Example: 250 GB input, 3 production workers

Live database object storage750 GB
Two full database backups1.5 TB
Calculated local disk per worker≈1.17 TB
RAM per worker256 GB
Total RAM across 3 workers768 GB
Total CPU across 3 workers96 vCPU

Object storage: 2.25 TB before content and growth. Local disk: 1.4 × 10 × 250 ÷ 3 ≈ 1,167 GB per worker; apply the release minimum if higher.

RAM and CPU use the documented production baseline of 256 GB and 32 vCPU per worker. They are not calculated from the 250 GB input volume. Query scope, concurrency, imports, and refresh jobs determine whether more capacity is needed.

In practice

Run it within your infrastructure and operating model.

Your cloud

Deploy in your cloud environment with supported object storage. Configure networking, identity, and access within your infrastructure.

On premises

Run on your Kubernetes infrastructure with a supported object store and worker-local disk.

Air-gapped

Use an internal image registry and controlled transfer of software, dependencies, and updates.

Your team operates the platform.

Customer administrators manage infrastructure, access, monitoring, backups, restores, and upgrades. DataWalk provides deployment guidance and administrator training as part of the handoff.

Installation establishes the environment. Ongoing operations keep it available and recoverable. Use-case implementation covers source mapping, business definitions, and analytical workflows. Assign owners for each.

Assign owners for each
1

Installation

Establishes the environment.

2

Ongoing operations

Keep it available and recoverable.

3

Use-case implementation

Covers source mapping, business definitions, and analytical workflows.

Environment prerequisites and compatibility

Prerequisites

  • An operational Kubernetes cluster and Helm administration skills.
  • At least 3 workers. The documented production baseline is 32 vCPU and 256 GB RAM per worker.
  • Supported object storage and local disks meeting capacity, throughput, and IOPS requirements.
  • Configured identity integration, certificates, networking, and access to deployment packages.

Confirm supported versions and disk requirements for the selected release. The minimum-requirements guide places workers in one Availability Zone. Define disaster recovery separately.

Development, test, and production

Each production instance requires at least one nonproduction environment; two are recommended. Test should match production. Development can be smaller.

Development, test, and production can use separate namespaces and Helm releases on one cluster. Allocate capacity, storage paths, ports, and monitoring for each.

The DevOps team must attend a deployment workshop. CKA or OpenShift administrator expertise is strongly recommended.

The supplied object-storage compatibility list includes Dell ECS, NetApp StorageGRID, Pure FlashBlade, and MinIO. Confirm support for the selected release.

Administrator handoff, monitoring, and upgrades

Training and installation

The handoff includes Deployment Certification and administrator training. Customer administrators must demonstrate installation and maintenance using DataWalk images and Helm charts.

Monitoring and recovery

Grafana and Prometheus provide monitoring. Configure alerts and backups to object storage. Test restores on separate infrastructure and disaster recovery against recovery point and recovery time objectives (RPO/RTO).

Upgrades

The documented procedure requires a database backup and Vault snapshot before replacing the Helm release with the Revive policy. Persistent data stays in object storage. Verify the service and recovery procedure after the upgrade.

Record the owners of infrastructure, backups, restores, and upgrades in the handoff. Define responsibilities for use-case implementation separately.

Resources

Review the evidence and the workloads.

Benchmark

Traversal benchmark

Test conditions, bounded and unbounded workloads, and complete request timings.

Read the benchmark (PDF)
Platform

Entity resolution

How matching produces resolved entities for use across the platform.

Read about entity resolution
Platform

Connected analytics

The graph, relational and search operations the architecture supports.

Review analytical methods
FAQ

Questions for the architecture review.

Is DataWalk a federated query layer?

No. Data for analysis is ingested and prepared in DataWalk. Existing warehouses and operational systems can remain in place as sources.

Does a shared state mean one server or one physical copy?

No. Compute runs across workers. The platform also uses local caches and backups, and sources and downstream exports remain separate.

How quickly do source changes reach an analysis?

After ingestion and the required dependent calculations finish. Set freshness targets for that full process, including downstream refreshes.

Does the 12 GB query budget describe a production deployment?

No. It is a query-memory budget from the traversal benchmark. Each test cluster had approximately 2.25 TiB of provisioned RAM. Production capacity depends on the workload and release requirements.

Review the architecture against your workload.

Bring representative queries, data volumes, freshness targets, and deployment constraints. Review capacity, operating responsibilities, and recovery requirements.

Book a demo

Architecture review checklist

Review inputs and outputs

Inputs
Initial load, growth forecast, content volume, representative queries, concurrency, refresh targets, and deployment constraints.
Tests
Query completeness and latency, cold and warm storage, ingestion overlap, permission changes, restores, and disaster recovery.
Decisions
Topology, versions, capacity, environments, RPO/RTO targets, and operating owners.
Outputs
A sizing and deployment plan, measured workload limits, and an administrator handoff checklist.