Articles

Giving Agents a Map

#research#agents#information retrieval#RAG#trustworthy AI

Beyond vector retrieval: a structured index can lift critical-fact coverage by more than a third.

Technical report PDF · 21 pages ↗

A Polished Report With Empty Spots

Ask an AI agent for a pre-investment due-diligence report and you often get something that looks quite respectable: company history, business lines, and market position, all neatly organized. Read it closely, however, and key information is missing:

board and ownership structure remain vague.

annual revenue and cash-flow evidence has large blanks.

core tendering details are again vague...

The report speaks fluently at the narrative level, then hesitates at the evidence boundary. This article examines how that omission happens, and how giving the agent a structured evidence mapopens another path for improvement.

Search Box vs. Case Directory: Two Retrieval Paradigms

The default way to answer over a private corpus is to give the agent vector retrieval; below, we call this the embedding agent[1]. The agent understands and rewrites a query, the system returns the top-k most similar documents, and the agent reads them before revising the query and asking again.

In that process, its grasp of the corpus structure, depth, and breadth is limited. Its view is bounded by retrieved results, so access to key information depends on repeated searches that try to exhaust the space. This engineering path works under one premise: the target evidence must be semantically close enough to the query wording.

Against this potentially black-box search process, we propose a new path based on trustworthy navigation:

Before interaction, build a structured index over the same corpus, then give the agent the ability to follow that index to evidence. The index works like a map: each module navigation page ultimately points to concrete documents. The agent that follows this directory, the index agent, acts more like an experienced analyst: reading the table of contents, judging branches, and opening the right documents.

the only thing that varies: how the agent reaches evidenceSame corpusthousands of docs per companyindex agentnavigates a structured index to evidenceembedding agenttop-k semantic search, round by roundSame recognizerfact-level gold annotations→ fact recall
How do we measure evidence coverage?

Study and benchmark design

Test corpus
Controlled variable: each company uses the exact same local document collection.
index agent
Experimental arm: navigates structured index paths to evidence documents.
embedding agent
Control arm: uses top-k semantic retrieval across multiple rounds.
Scoring
Same evaluator, measuring fact recall against fact-level gold annotations.
Only variable
The retrieval mechanism by which the agent reaches and obtains evidence.

The benchmark contains 50 due-diligence questions: 10 companies, each with 5 core question types. We use two scoring views:

near-hiteffective access

Partial coverage of the key fact, used to test whether the agent reached the right evidence source.

strictstrict coverage

The key fact is fully and accurately extracted and stated.

Note: all body results include 95% company-cluster bootstrap confidence intervals, with significance from exact company-cluster permutation tests; full protocol in the paper.

The Coverage Lift From Giving the Agent a Map

Across all four primary metrics the index agent leads, with planned exact company-cluster permutation p values of at most 0.006. The largest gap is where it matters most for due diligence: coverage of critical facts.

critical near-hit p 0.006 · 95% CIembedding0.400index0.550critical strict p 0.004 · 95% CIembedding0.130index0.269overall near-hit p 0.002 · 95% CIembedding0.270index0.405overall strict p 0.002 · 95% CIembedding0.083index0.195

Critical counts only the high-value key-fact subset; overall counts all scored target facts (critical, important, and relevant). Error bars are 95% company-cluster bootstrapCIs; p values are exact company-cluster permutation tests.

Seeing the Path: Evidence Access Becomes Auditable

Below is the actual Case A run — every document, every search, every step of the index agent's walk, projected onto stacked recall-state planes over a shared 2D document embedding.[2] The visualization is not only a demo: it turns evidence access from black-box similarity matching into an auditable action trace.

loading manifold…

Long-Tail and Hidden Evidence Retrieval

Top-k retrieval is a ranking contest that restarts every round. Overview-like documents naturally match many common keywords and keep sitting near the front; the hard facts that decide the answer often hide in single-point documents with only a weak semantic overlap to the query.

In our tests, three kinds of key evidence sink especially often:

event-level evidence

Case A — investment due diligence

critical near-hit 0.333 vs 0.750

A key financing event and a regulatory-list news item were both sitting quietly in the corpus, but never surfaced across the embedding agent's six searches. The index agent followed financials/financing_history and risk directly to them, and its answer covered the $1B competitive-bid financing, 30× oversubscription, and CMC and FCC listings.

numeric field evidence

Case B — credit / supplier access

critical near-hit 0.111 vs 0.667

Credit review depends on structured financial fields. The embedding agent searched in the right direction — slowing growth, three years of negative operating cash flow — but the key fields were almost empty in the retrieved material. The index agent opened one prospectus-analysis document under financials/ and read the table out: three-year revenue of 664M, 755M, and 820M RMB, cumulative losses of 1.59B RMB, and a 96.93% subscription-revenue share.

negative-turn evidencefixed-budget

Case C — industry & sales analysis

critical near-hit 0.273 vs 0.909

Companies rarely narrate their own pain points, so negative turning points sink deep in a similarity space dominated by official material. The index agent's high-signal module surfaced the flagship-project failure, the 200-person layoff plan, and the ranking drop to #27, then turned them into concrete B2B sales entry points. This question comes from the fixed-budget evaluation described below: both agents faced the same $0.40 model-API budget threshold during evidence access, and this is the largest per-question gap in that evaluation.

The sinking pattern is especially visible in negative-turn evidence. Company C is a B2C game company; when the task is to find actionable B2B sales entry points,[3] the difference between the two retrieval designs gradually appears. Start with the search process:

embedding agent: six rewritten queries

  • {company} business positioning company profile main business game development and publishing
  • {company} games product matrix main game products
  • {company} network technology hiring financing tendering investment events
  • {company} games tenders government procurement winning bids bid information cooperation projects
  • {company} games competition peers industry competitive landscape market ranking
  • {company} product updates new game launches version updates 2024 2025 2026

index agent: read the directory layer, then drill into evidence regions

  • read index://overview: sees how the corpus is organized
  • open region://business-and-transactions: looks for cooperation, procurement, and IP entry points
  • open region://public-events: locates copyright disputes and enforcement reports
  • open region://risk-and-counterparty: checks supplier anomalies, compliance, and blacklist signals
  • open evidence://candidate-docs: grounds the leads in citable text
  • synthesize answer: turns evidence into sales-entry points and risk judgment

The embedding agent's six queries mostly describe the company's positive profile. The only B2B-adjacent round is narrowed to tenders, government procurement, and winning bids, so it keeps retrieving official pages, encyclopedic summaries, hiring, product releases, and rankings. The right side shows a de-identified read path: it preserves the evidence-access logic while hiding real directories, file hierarchy, and internal index names.

Different paths, different answers:

embedding agent: three “not covered” statements

Customer leads: Core conclusion: Company C is a typical B2C game company;B2B customer leads are not covered in fetched evidence.

B-side cooperation leads (limited): Company C is itself an investor and has invested in several game companies; portfolio companies may form a business ecosystem, with public contact channels visible.

Gap note:No concrete customer list, signing record, or cooperation case was found for Company C as a buyer or supplier of B2B products or services.

index agent: three leads

B2B supply-chain leads:Advertising/material supplier: a video platform sued Company C for copyright infringement over advertising materials;anti-cheat demand: a core product plug-in case and plug-in proliferation in a new product;internal anti-corruption involving suppliers: 70+ violations investigated in 2024-2025, with 22 external suppliers blacklisted.

Sales-entry analysis: technical services can enter through anti-cheat/security defense, cloud/server operations, and UE engine support; compliance or audit tools can enter through supplier-risk systems and internal-audit SaaS.

Semantic retrieval can only hit what you already thought to ask for; a structured directory can surface evidence you did not anticipate through an auditable path. Advertising-material disputes, anti-cheat cases, and similar events leave textual traces, but rewriting the query to hit them is closer to a lucky contest.

The Boundary of the Map: Not a Silver Bullet

Although the index agent wins overall, it does not win every question. By per-question overall recall, it shows limitations in 3 of the 50 questions.[4]

largest overall-recall reversalCase D: industry and sales analysis

The target answer is concentrated in a few fact-dense industry articles and one procurement result, rather than scattered across deep typed index leaves.

Where the embedding agent landed

final used: 10 docs

  • Search 5 customers / hospital channels / sales network / distributors / tenders / medical procurement
  • Search 6 competitors / industry structure / volume procurement impact / innovative-drug globalization
Key final sources
  • “innovative-drug approvals and BD acceleration”
  • “globalization transition”
  • “preclinical pharmacology and patent-layout consulting procurement result”

Where the index agent drifted

final candidate-only: 8 docs

  • It opened customer, supplier, competitor, financial, and risk nodes.
  • Its retained final candidates leaned toward patents, hiring posts, and bid results, so the industry-article sales framing did not fully enter the answer.
Key final candidates
  • ADC preparation patent
  • siRNA / QC / trainee hiring
  • isolator line, stopper washer, and bioreactor bid results

We think the boundary is clear: when a task needs fine-grained evidence from discrete, long-tail nodes, the structured map helps; when the answer is naturally concentrated in a few dense documents near the corpus center, traditional semantic retrieval can already be efficient enough.

Robustness and Cost

Beyond the main result, we test whether the gain comes from navigation structure, whether it remains observable under a shared evidence-access budget constraint, and what the interface costs.

Ablation: The Contribution of Navigation Structure

Knowing which path performs better is not enough; we also want to know where the gain comes from. We ran an ablation study that stripped the index environment layer by layer: first removing skill-prompt text, then removing navigation documents, and finally flattening it into a structure-free base environment to see how critical recall changes.

critical near-hit pooled · 95% CIfull index agent0.558no skill text0.551no navigation0.418flat base env0.382

Pooled critical near-hit (95% company-cluster bootstrap CI). Bars, top to bottom: full index agent, skill text removed, navigation documents removed, flat base environment.

Removing skill text costs almost nothing (−0.007); removing navigation documents collapses recall (−0.133); the flat base environment (0.382) lands at the embedding agent's level (0.400). The gain is in structure, not in the prompt.

Performance Under a Shared Budget Constraint

The natural objection is whether the index agent performs better simply because it can spend more. To check this boundary, we ran a separate fixed-budget evaluation. Each agent first used its native interface to search or navigate evidence. Once the threshold was reached, tool use stopped and the agent synthesized its final answer from the evidence and interaction history already available.

overall near-hit (fixed-budget) p 0.059 · 95% CIembedding0.254index0.297

Overall near-hit (95% company-cluster bootstrap CI). 48 of 50 embedding runs and 50 of 50 index runs reached the stopping threshold. The shared threshold constrains both interfaces but does not equalize realized spending, retrieval counts, or tool granularity; full protocol in the paper.

At the shared $0.40 budget point, index records an overall near-hit of 0.297 versus embedding's 0.254 (p = 0.059) — the positive signal is visible but does not reach significance. This is an operating-boundary observation, not an estimate of how much of the main-study gap is attributable to spending.

Resource Cost

The index agent used about 1.26M tokens per run, while the embedding agent used only 198K: more than a sixfold difference. Its candidate surface is also much wider — roughly 399 candidates per question versus 83 — yet opened evidence does not grow proportionally, and the final answer does not always carry the opened evidence through safely.

Limitations and Ongoing Work

Making terrain visible is what this map already does. Making the agent finish the right terrain before answering is the next step: route review, pre-synthesis coverage checks, and budget-aware stopping are all in progress.

  • We study company due diligence; transfer to other domains remains to be validated.
  • The experimental agent is claude code with deepseek-v4-flash; stability across models and other interaction variables remains future work.
  • Navigation-structure design is still being optimized, and index-series comparison experiments are underway.

For the complete experimental design, evaluation contract, statistical results, and appendix, see the technical report (PDF, 21 pages).


  1. Vector retrieval as a default starting point is visible in mainstream documentation: OpenAI file search / Retrieval API performs semantic retrieval over vector stores, and Anthropic's engineering guidance notes that many AI-native applications use embedding-based retrieval before reasoning. Production systems often add hybrid dense + BM25 retrieval, agentic retrieval, or filesystem-style navigation, but vector retrieval remains the common starting point for private document QA.
  2. Layout: documents are a 2D PCA of the company's full evidence-document embeddings, repeated across recall-state planes; navigation and query anchors are embedded with the same model and projected into the same basis. Distances are approximate projections — exact similarity lives in the original embedding space.
  3. Answer and trace excerpts are translated from Chinese and de-identified; scores, counts, and document paths are verbatim from the scoring artifacts and case files.
  4. Per-question overall-recall reversals are computed from paired per-question differences in the scored comparison summary; the critical near-hit reversal count is reported as mechanism context.