See why the records belong together.
Resolve people and organizations inside DataWalk. Follow the activity that matching brings together, with the source records and evidence still available to inspect.
The wrong identity changes the answer.
Split one person across records and their activity appears unrelated. Join two people by mistake and a connection appears where none exists.
A matching result needs more than a score. Your team needs to see which evidence supports it and what conflicts with it.
Your data isn’t all in tidy rows.
PDFs, Word documents, emails, spreadsheets. More files arrive every day. DataWalk brings these sources together and extracts people, organizations and relationships from document text.
- Word
- Spreadsheet
The evidence stays with the match.
Resolve identities where you analyze them. Source records, matching evidence and resolution links remain queryable together, without a separate matching store to reconcile.
Compare more than the spelling.
Map source data to the supported people and organization ontology. DataWalk normalizes values and weighs evidence across names, addresses, identifiers and other configured signals.
A shared switchboard number should not carry the weight of a distinctive identifier.
The model defines what a record represents; resolution evaluates which records refer to the same entity. A missed match may mean a pair was never compared, or that its evidence did not support a link.
Inspect what made the match.
See which records were compared, which values agreed and which rule changed the verdict. Matching names can point one way while conflicting identifiers point another.
The evidence remains queryable alongside the resolution results. Some records remain unresolved or need further review; possible-match evidence is queryable, but a dedicated review workflow must be scoped separately.
Same name does not settle the match
| Evidence | Record A | Record B | Signal |
|---|---|---|---|
| Name | Morgan Rivera | Morgan Rivera | Agrees |
| Phone | Shared office | Shared office | Agrees · low weight |
| Exclusive ID | ID-104 | ID-908 | Conflicts |
The identifiers conflict
Under a policy that treats these IDs as exclusive, the conflict can override the name agreement.
Illustrative evidence, not a measured match result. Your configuration determines the verdict.
Follow the identity across sources.
A customer record and an external registry entry can stay separate while an identity link connects the relationships each holds.
Search, analysis and applications can follow those identity links. A new match can expose activity that no individual source could connect.
162 million organization records.
3.69 billion candidate pairs evaluated in 7h 03m on three computation nodes, or 3h 04m on twelve.
These are active compute times. Result import is measured separately. Four times the nodes produced approximately 2.30 times the active-compute speed on this workload. Evaluate false matches and missed matches separately on labeled examples from your data.
162.07M organization records · 3.69B candidate pairs. Asynchronous result import excluded. Measured workload; active compute only.
Resolve a corpus and the changes that follow.
Initial batch
Test known matches and known non-matches from your sources before running the corpus. Inspect the evidence behind both, then run the batch with the configuration you have evaluated.
Output: Corpus links and evidence.
Incremental updates
After the initial build, incremental processing identifies changed records and the comparisons they affect. Update the identity picture without rerunning the full batch for every change.
Output: Affected links and evidence.
On-demand check
On-demand scoring compares a submitted record with the already-processed corpus. It returns matching evidence without writing that submitted record into the knowledge graph.
Output: Response against the processed corpus.
Reference-corpus freshness matters; verify configuration compatibility after changes.
Evaluate resolution on your data.
Prepare an evaluation
Choose sample cases and tests for your data.
Enterprise Knowledge Graph
See how resolved identities connect to the rest of the business model.
Connected Analytics
Follow the relationships those identities bring into view.
Questions about matching.
What can DataWalk resolve?
Native entity resolution supports people and organizations. Addresses, phones, identifiers and other supported values provide matching evidence; they are not all resolved endpoint types.
Are source records merged automatically?
Resolution produces links and clusters by default. Merging is optional and reversible. Your team chooses how to use the results. Agree how identity changes affect current analyses, retained results and downstream copies.
Can we inspect why a pair did not match?
The pipeline exposes candidate-generation and scoring evidence. Investigating a missing match may mean checking whether the records were compared at all. Non-match links are not populated by default.
Is there a review queue for possible matches?
Possible-match results and evidence are available to query. DataWalk ER does not include a dedicated review workflow. Review and operational handling should be scoped as part of the implementation.
How should we evaluate name matching?
Name matching is tuned primarily for Western naming conventions. Test the languages, scripts and naming patterns in your data, including the difficult matches and the records that must stay separate.
Is on-demand scoring a real-time identity service?
It scores a request against the processed corpus. The documented tests report responses in seconds, not a sub-second service-level commitment or a concurrent-throughput guarantee.
Test the identities your analysis depends on.
Bring the records that should match and the ones that must stay apart. Inspect the evidence for both.
Plan an ER evaluation
- Bring
- Known matches, known distinct entities, uncertain cases, common names, shared contacts, conflicting identifiers and relevant languages/scripts.
- Try
- Introduce contradictory evidence; inspect the supported resolution update, rerun an affected analysis, and check retained/exported results separately.
- Check
- Report false joins, missed matches, unresolved cases and candidate-generation exclusions separately, alongside time and review effort. Verify on-demand configuration compatibility.
- Review output
- An error breakdown and configured-policy findings.