Ask an AI agent for a pre-investment due-diligence report and you often get something that looks quite respectable: company history, business lines, and market position, all neatly organized. Read it closely, however, and key information is missing:
board and ownership structure remain vague.
annual revenue and cash-flow evidence has large blanks.
core tendering details are again vague...
The report speaks fluently at the narrative level, then hesitates at the evidence boundary. This article examines how that omission happens, and how giving the agent a structured evidence mapopens another path for improvement.
Search Box vs. Case Directory: Two Retrieval Paradigms
The default way to answer over a private corpus is to give the agent vector retrieval; below, we call this the embedding agent[1]. The agent understands and rewrites a query, the system returns the top-k most similar documents, and the agent reads them before revising the query and asking again.
In that process, its grasp of the corpus structure, depth, and breadth is limited. Its view is bounded by retrieved results, so access to key information depends on repeated searches that try to exhaust the space. This engineering path works under one premise: the target evidence must be semantically close enough to the query wording.
Against this potentially black-box search process, we propose a new path based on trustworthy navigation:
Before interaction, build a structured index over the same corpus, then give the agent the ability to follow that index to evidence. The index works like a map: each module navigation page ultimately points to concrete documents. The agent that follows this directory, the index agent, acts more like an experienced analyst: reading the table of contents, judging branches, and opening the right documents.
How do we measure evidence coverage?
Study and benchmark design
Test corpus
Controlled variable: each company uses the exact same local document collection.
index agent
Experimental arm: navigates structured index paths to evidence documents.
embedding agent
Control arm: uses top-k semantic retrieval across multiple rounds.
Scoring
Same evaluator, measuring fact recall against fact-level gold annotations.
Only variable
The retrieval mechanism by which the agent reaches and obtains evidence.
The benchmark contains 50 due-diligence questions: 10 companies, each with 5 core question types. We use two scoring views:
near-hiteffective access
Partial coverage of the key fact, used to test whether the agent reached the right evidence source.
strictstrict coverage
The key fact is fully and accurately extracted and stated.
Note: all body results include 95% company-cluster bootstrap confidence intervals, with significance from exact company-cluster permutation tests; full protocol in the paper.
The Coverage Lift From Giving the Agent a Map
Across all four primary metrics the index agent leads, with planned exact company-cluster permutation p values of at most 0.006. The largest gap is where it matters most for due diligence: coverage of critical facts.
Critical counts only the high-value key-fact subset; overall counts all scored target facts (critical, important, and relevant). Error bars are 95% company-cluster bootstrapCIs; p values are exact company-cluster permutation tests.
Seeing the Path: Evidence Access Becomes Auditable
Below is the actual Case A run — every document, every search, every step of the index agent's walk, projected onto stacked recall-state planes over a shared 2D document embedding.[2] The visualization is not only a demo: it turns evidence access from black-box similarity matching into an auditable action trace.
loading manifold…
Long-Tail and Hidden Evidence Retrieval
Top-k retrieval is a ranking contest that restarts every round. Overview-like documents naturally match many common keywords and keep sitting near the front; the hard facts that decide the answer often hide in single-point documents with only a weak semantic overlap to the query.
In our tests, three kinds of key evidence sink especially often:
event-level evidence
Case A — investment due diligence
critical near-hit 0.333 vs 0.750
A key financing event and a regulatory-list news item were both sitting quietly in the corpus, but never surfaced across the embedding agent's six searches. The index agent followed financials/financing_history and risk directly to them, and its answer covered the $1B competitive-bid financing, 30× oversubscription, and CMC and FCC listings.
numeric field evidence
Case B — credit / supplier access
critical near-hit 0.111 vs 0.667
Credit review depends on structured financial fields. The embedding agent searched in the right direction — slowing growth, three years of negative operating cash flow — but the key fields were almost empty in the retrieved material. The index agent opened one prospectus-analysis document under financials/ and read the table out: three-year revenue of 664M, 755M, and 820M RMB, cumulative losses of 1.59B RMB, and a 96.93% subscription-revenue share.
negative-turn evidencefixed-budget
Case C — industry & sales analysis
critical near-hit 0.273 vs 0.909
Companies rarely narrate their own pain points, so negative turning points sink deep in a similarity space dominated by official material. The index agent's high-signal module surfaced the flagship-project failure, the 200-person layoff plan, and the ranking drop to #27, then turned them into concrete B2B sales entry points. This question comes from the fixed-budget evaluation described below: both agents faced the same $0.40 model-API budget threshold during evidence access, and this is the largest per-question gap in that evaluation.
The sinking pattern is especially visible in negative-turn evidence. Company C is a B2C game company; when the task is to find actionable B2B sales entry points,[3] the difference between the two retrieval designs gradually appears. Start with the search process:
embedding agent: six rewritten queries
{company} business positioning company profile main business game development and publishing
{company} games tenders government procurement winning bids bid information cooperation projects
{company} games competition peers industry competitive landscape market ranking
{company} product updates new game launches version updates 2024 2025 2026
index agent: read the directory layer, then drill into evidence regions
read index://overview: sees how the corpus is organized
open region://business-and-transactions: looks for cooperation, procurement, and IP entry points
open region://public-events: locates copyright disputes and enforcement reports
open region://risk-and-counterparty: checks supplier anomalies, compliance, and blacklist signals
open evidence://candidate-docs: grounds the leads in citable text
synthesize answer: turns evidence into sales-entry points and risk judgment
The embedding agent's six queries mostly describe the company's positive profile. The only B2B-adjacent round is narrowed to tenders, government procurement, and winning bids, so it keeps retrieving official pages, encyclopedic summaries, hiring, product releases, and rankings. The right side shows a de-identified read path: it preserves the evidence-access logic while hiding real directories, file hierarchy, and internal index names.
Different paths, different answers:
embedding agent: three “not covered” statements
Customer leads: Core conclusion: Company C is a typical B2C game company;B2B customer leads are not covered in fetched evidence.
B-side cooperation leads (limited): Company C is itself an investor and has invested in several game companies; portfolio companies may form a business ecosystem, with public contact channels visible.
Gap note:No concrete customer list, signing record, or cooperation case was found for Company C as a buyer or supplier of B2B products or services.
index agent: three leads
B2B supply-chain leads:Advertising/material supplier: a video platform sued Company C for copyright infringement over advertising materials;anti-cheat demand: a core product plug-in case and plug-in proliferation in a new product;internal anti-corruption involving suppliers: 70+ violations investigated in 2024-2025, with 22 external suppliers blacklisted.
Sales-entry analysis: technical services can enter through anti-cheat/security defense, cloud/server operations, and UE engine support; compliance or audit tools can enter through supplier-risk systems and internal-audit SaaS.
Semantic retrieval can only hit what you already thought to ask for; a structured directory can surface evidence you did not anticipate through an auditable path. Advertising-material disputes, anti-cheat cases, and similar events leave textual traces, but rewriting the query to hit them is closer to a lucky contest.
The Boundary of the Map: Not a Silver Bullet
Although the index agent wins overall, it does not win every question. By per-question overall recall, it shows limitations in 3 of the 50 questions.[4]
largest overall-recall reversalCase D: industry and sales analysis
The target answer is concentrated in a few fact-dense industry articles and one procurement result, rather than scattered across deep typed index leaves.
“preclinical pharmacology and patent-layout consulting procurement result”
Where the index agent drifted
final candidate-only: 8 docs
It opened customer, supplier, competitor, financial, and risk nodes.
Its retained final candidates leaned toward patents, hiring posts, and bid results, so the industry-article sales framing did not fully enter the answer.
Key final candidates
ADC preparation patent
siRNA / QC / trainee hiring
isolator line, stopper washer, and bioreactor bid results
We think the boundary is clear: when a task needs fine-grained evidence from discrete, long-tail nodes, the structured map helps; when the answer is naturally concentrated in a few dense documents near the corpus center, traditional semantic retrieval can already be efficient enough.
Robustness and Cost
Beyond the main result, we test whether the gain comes from navigation structure, whether it remains observable under a shared evidence-access budget constraint, and what the interface costs.
Ablation: The Contribution of Navigation Structure
Knowing which path performs better is not enough; we also want to know where the gain comes from. We ran an ablation study that stripped the index environment layer by layer: first removing skill-prompt text, then removing navigation documents, and finally flattening it into a structure-free base environment to see how critical recall changes.
Pooled critical near-hit (95% company-cluster bootstrap CI). Bars, top to bottom: full index agent, skill text removed, navigation documents removed, flat base environment.
Removing skill text costs almost nothing (−0.007); removing navigation documents collapses recall (−0.133); the flat base environment (0.382) lands at the embedding agent's level (0.400). The gain is in structure, not in the prompt.
Performance Under a Shared Budget Constraint
The natural objection is whether the index agent performs better simply because it can spend more. To check this boundary, we ran a separate fixed-budget evaluation. Each agent first used its native interface to search or navigate evidence. Once the threshold was reached, tool use stopped and the agent synthesized its final answer from the evidence and interaction history already available.
Overall near-hit (95% company-cluster bootstrap CI). 48 of 50 embedding runs and 50 of 50 index runs reached the stopping threshold. The shared threshold constrains both interfaces but does not equalize realized spending, retrieval counts, or tool granularity; full protocol in the paper.
At the shared $0.40 budget point, index records an overall near-hit of 0.297 versus embedding's 0.254 (p = 0.059) — the positive signal is visible but does not reach significance. This is an operating-boundary observation, not an estimate of how much of the main-study gap is attributable to spending.
Resource Cost
The index agent used about 1.26M tokens per run, while the embedding agent used only 198K: more than a sixfold difference. Its candidate surface is also much wider — roughly 399 candidates per question versus 83 — yet opened evidence does not grow proportionally, and the final answer does not always carry the opened evidence through safely.
Limitations and Ongoing Work
Making terrain visible is what this map already does. Making the agent finish the right terrain before answering is the next step: route review, pre-synthesis coverage checks, and budget-aware stopping are all in progress.
We study company due diligence; transfer to other domains remains to be validated.
The experimental agent is claude code with deepseek-v4-flash; stability across models and other interaction variables remains future work.
Navigation-structure design is still being optimized, and index-series comparison experiments are underway.
For the complete experimental design, evaluation contract, statistical results, and appendix, see the technical report (PDF, 21 pages).
Vector retrieval as a default starting point is visible in mainstream documentation: OpenAI file search / Retrieval API performs semantic retrieval over vector stores, and Anthropic's engineering guidance notes that many AI-native applications use embedding-based retrieval before reasoning. Production systems often add hybrid dense + BM25 retrieval, agentic retrieval, or filesystem-style navigation, but vector retrieval remains the common starting point for private document QA. ↩
Layout: documents are a 2D PCA of the company's full evidence-document embeddings, repeated across recall-state planes; navigation and query anchors are embedded with the same model and projected into the same basis. Distances are approximate projections — exact similarity lives in the original embedding space. ↩
Answer and trace excerpts are translated from Chinese and de-identified; scores, counts, and document paths are verbatim from the scoring artifacts and case files. ↩
Per-question overall-recall reversals are computed from paired per-question differences in the scored comparison summary; the critical near-hit reversal count is reported as mechanism context. ↩