Skip to content

HDF MCP Server ​

hdf mcp runs the Heimdall Data Format server over the Model Context Protocol (MCP), so an AI agent or MCP-aware client can read, analyze, and author HDF documents through a small, typed tool surface instead of shelling out to the CLI. It speaks JSON-RPC over stdio: the client launches hdf mcp as a subprocess and every request and response is a JSON-RPC frame on stdin/stdout (nothing else is written to stdout).

bash
hdf mcp        # start the server (stdio transport)

A typical MCP client entry (for example a config.toml mcp_servers block) launches the built binary with mcp as its argument and sets the environment variables below.

This guide documents the shipped surface.

The tool surface ​

The server exposes ten tools. Read tools never mutate anything; write tools are governed by a deployer-controlled gate (see Reads vs. writes). Every tool returns a compact summary plus a reusable handle — never a multi-megabyte document body.

Read tools ​

ToolPurpose
hdf_openEntry point: open a document, return its detected type, schema version, validity, and a handle. Optional — every read tool also accepts a {path} directly.
hdf_inspectDocument structure and metadata for all eight HDF document types (counts, component/baseline/assessment breakdowns, top-level fields). For a results document the metadata includes the root tool (name, version, format — each when present) and generator when present, and each baseline entry carries its labels when it has any — on a multi-baseline document that is how an agent sees which baseline is which. It never returns requirements.
hdf_queryThe only path to requirements (results and baseline documents only). Filters by status, severity, requirement ID, tag, and free text; concise by default, full on request; paginates when a response would exceed the token budget. An opt-in fields array adds cross-source correlation keys per row (see Correlation fields). Takes one source, or several results documents as one set via sources[] (see Several documents as one set).
hdf_complianceStatus × severity rollups, the compliance percentage, optional threshold verdicts, and the agent-attributed override count for one document — or for a set of results documents via sources[]. groupBy is baseline (one group per input baseline), severity, nistFamily, tool (the scanner each baseline came from — the label the combined view carries; a single document groups as unlabeled) or cwe (each requirement's CWE numbers, unmapped for none).
hdf_aggregateRoll up status/severity counts across multiple documents in one call — per-source counts plus a server-computed total and compliance, with optional status/severity/nist filters. Counts only, never rows; a source that fails to load is reported and the rest still aggregate.
hdf_diffCompare two documents — temporal (results across time) or system-drift (system documents) — and emit an hdf-comparison.
hdf_validateValidate a document in schema, checksums, or completeness mode.

The hdf_inspect vs. hdf_query bright line is deliberate: inspect is structure, query is requirements. If you want to know how many baselines a results file has, or what components a system document declares, call hdf_inspect. If you want the requirements themselves — their statuses, severities, or IDs — call hdf_query. hdf_inspect will never return requirements, and hdf_query only accepts results and baseline documents.

Write tools ​

ToolPurpose
hdf_convertConvert source security-tool output (Nessus, SARIF, gosec, VEX, …) into an HDF document. Auto-detects the source format; returns a summary and a handle. Enforces the same requirement-count fidelity check as hdf convert (see Conversion Count Fidelity): a conversion that lost findings is refused with SCHEMA_INVALID.
hdf_authorAuthor an HDF document from model-supplied structured content: system (components), plan (assessments), evidence (contents), or amendments (overrides). For amendments the server holds field authority (see below).
hdf_apply_amendmentApply an hdf-amendments document to an hdf-results document, producing a new results file with effectiveStatus/effectiveImpact/disposition computed and the before/after compliance delta reported. It never overwrites its results input.

hdf_author is a single authoring tool. It subsumes what earlier designs split into separate document-builder and amendment-creator tools: the model supplies the content array, and the server assembles, validates, and stamps the schema-valid envelope.

The source model ​

Every tool that reads a document takes a source:

json
{ "source": { "path": "results/rhel9.json" } }

or

json
{ "source": { "handle": "eyJwYXRoIjoi…" } }
  • A {path} is resolved under HDF_MCP_ROOT (see Deployment); a path that escapes the root is refused.
  • A {handle} is the self-describing identity a prior call returned. It encodes the document's path, content hash, detected type, and schema version, so the server can re-read it and detect if the file changed underneath (a stale handle is reported, not silently trusted).

hdf_open is optional, not a mandatory first hop — you can pass a {path} straight to hdf_inspect, hdf_query, or any read tool. Its value is minting a handle you then reuse across a multi-step workflow.

A document a write tool produces is held in the same cache, so its handle resolves on the next call. A document larger than HDF_MCP_CACHE_BYTES cannot be, and rather than return a handle that will fail later, hdf_convert, hdf_author and hdf_apply_amendment attach a notice saying so and telling you to set output and pass that path instead. In batch conversion the notice is per entry, since one oversized file does not affect the rest.

Agents pass handles back to tools, never document bodies. Responses are summaries plus handles precisely so a large results file never has to travel through the model's context to be operated on again.

Several documents as one set ​

hdf_query and hdf_compliance also take sources[] — several results documents that the server combines in memory, for that call into one multi-baseline view (the engine Merge). Nothing is written and no merged file exists: the pipeline keeps one document per scanner, which is what thresholds, amendments and exports work on, and the agent asks its cross-tool question over any set of them.

json
{
  "sources": [
    { "path": "scans/gosec.hdf.json" },
    { "path": "scans/zap.hdf.json" },
    { "path": "scans/grype.hdf.json" }
  ],
  "baseline": "owasp zap/*",
  "fields": ["cwe"],
  "limit": 2
}

Every baseline in the view is renamed <tool>/<original name> (the tool is the document's root tool.name) and labelled tool / toolVersion / sourceDocument, so a baseline glob selects one scanner across the set, full rows say which scanner each row came from, and in hdf_compliancegroupBy: baseline is one group per input baseline with its baselineIndex, groupBy: tool is one rollup per scanner (keyed by that label, never by the name prefix), and groupBy: cwe rolls the set up by weakness across scanners — by each requirement's normalized cwe[] field, the same key hdf_query's fields: ["cwe"] reads. A converter that records CWEs only under tags.cwe (the SARIF converter does today) leaves that field empty, and its rows group as unmapped until it is fixed. A threshold is evaluated over the set's counts exactly as over one document.

The response names the set instead of one handle — each member by the source it was loaded from, not by a handle (a handle per member cost ~230 bytes on every response for something you already hold; open a member with hdf_open if you want one):

json
{ "sources": [ { "index": 0, "source": "scans/gosec.hdf.json" },
               { "index": 1, "source": "scans/zap.hdf.json" },
               { "index": 2, "source": "scans/grype.hdf.json" } ],
  "docType": "results", "total": 28, "returned": 2, "truncated": true, "requirements": [ … ] }

Rules: pass exactly one of source / sources; a one-element sources is the single-source call, unchanged; every member must be a valid results document (a baseline document can be queried alone, not in a set), and a member that fails to load, is another type, or is schema-invalid refuses the whole set naming its index (sources[1] is a system document …) — a set is never partially answered. If two baselines still share a name after the tool prefix, notice says so; rows and groups are keyed by position, so the answer is unaffected.

Reads vs. writes ​

The server draws a firm line between reading and writing.

Reads degrade gracefully. A structurally-recognizable but schema-invalid document still opens: the read returns valid: false and the best structural summary it can, rather than failing. This lets an agent inspect and reason about imperfect documents.

Writes refuse invalid input. Every write tool validates the document it would produce against the real schema before writing, and refuses a document that does not validate (SCHEMA_INVALID) — it is never written or handed back.

Write tools write by default; dry_run previews. Within a deployment that permits writes, a write tool writes its output file by default; passing dryRun: true returns the same summary and validation verdict but writes nothing.

HDF_MCP_ENABLE_WRITES is the deployer ceiling. Writes are disabled by default. When they are disabled, a write call still succeeds — it returns the computed summary plus a WRITES_DISABLED notice — but touches no file. The agent cannot lift this ceiling; only the deployer can, by setting the variable.

Agent attribution travels in the data. When hdf_author drafts amendments from model judgment, the server stamps each override with appliedBy.type = "agent" and an appliedAt timestamp — the model cannot claim a non-agent identity. Those overrides survive hdf_apply_amendment onto the applied results file, so hdf_compliance (and the hdf validate / hdf evidence verify CLI readouts) can report how many overrides an agent applied, without back-tracing to the amendments document. This detective surface is what makes autonomous authoring auditable.

Deployment ​

The server is configured through environment variables (tool selection also has a matching --tools flag):

VariablePurposeDefault
HDF_MCP_ROOTPath-confinement root. Every source.path and write output is resolved under it; anything resolving outside is refused (PATH_DENIED).the process working directory
HDF_MCP_ENABLE_WRITESThe write gate. Truthy (1, true, yes, on) permits writes; anything else disables them (write calls return previews).disabled
HDF_MCP_CACHE_BYTESBudget for the byte-bounded LRU parsed-document cache, so repeated reads of the same document skip re-parsing.256 MB
HDF_MCP_MAX_SIZEPer-document input ceiling in bytes. A file larger than this is refused (TOO_LARGE) before it is read into memory, and the same ceiling applies to a parsed document.50 MB
HDF_MCP_LOG_LEVELStructured-log level written to stderr (error, warn, info, debug). stdout carries only JSON-RPC frames.info
HDF_MCP_TOOLSWhich tools to advertise: a comma-separated list of tool names (hdf_open,hdf_query) or a profile — read (the read/analysis tools) or all. Advertising fewer tools shrinks the schema an agent re-sends every turn. The --tools flag overrides it.all tools

Set HDF_MCP_ROOT to the directory that holds the documents the agent should work with, and leave HDF_MCP_ENABLE_WRITES unset for a read-only trial — the write tools still work, returning previews, so you can see exactly what they would produce before granting write access.

Advertising fewer tools ​

Every request an agent makes re-sends the whole tool surface, so the set of tools you advertise is a fixed per-turn cost. A deployment that only reads HDF documents does not need the hdf_convert, hdf_author, or hdf_apply_amendment write tools on every turn — and an agent built for one job may need only a couple of tools.

Select the surface with --tools (or HDF_MCP_TOOLS), taking either explicit tool names or a profile:

bash
hdf mcp --tools read                 # the read/analysis tools only
hdf mcp --tools hdf_open,hdf_query   # a focused two-tool surface
hdf mcp                              # every tool (the default)

The read profile is hdf_open, hdf_inspect, hdf_query, hdf_compliance, hdf_aggregate, hdf_diff, and hdf_validate. Measured in the model-facing representation — the tool name, description, and parameters a provider actually sends the model each turn — the full ten-tool surface is ~3,116 tokens; the read profile is ~1,877 (a 40% cut), and a two-tool hdf_open,hdf_query surface is ~573 (82%). An unknown tool name is rejected at startup rather than silently dropped, so a bad launch config fails loudly.

Recommended for an unattended analysis pipeline: --tools read. A CI/CD harness that reads and rolls up already-converted HDF documents — the token-sensitive case, running many turns with a human only at the end — needs none of the convert/author/apply_amendment write tools, and paying for them on every turn is pure overhead. Set --tools read (add hdf_convert if the pipeline also converts raw scanner output) and the surface drops ~40%.

The default is deliberately every tool, not read. The server cannot know whether a given deployment only analyzes or also converts and authors, so the safe default is capability-complete; narrowing is the operator's explicit, non-breaking choice. (Defaulting to read would silently drop the write tools from an existing deployment that relied on them — a regression traded for a config line each analysis deployment can set itself.)

What the read surface deliberately does not return ​

Conversion is close to lossless: hdf convert preserves each scanner's original finding verbatim in the requirement's code field. No read tool projects code, at any verbosity — a deliberate boundary rather than an oversight.

This is not about breaching a token bound. The per-response budgets (concise 2,000, full 10,000) are structural: hdf_query packs rows until the budget is reached and paginates the rest, so a fatter row simply yields fewer rows per page — no field can push a response over its bound. Measured on a real Grype scan of 89 findings, concise returns ~48 rows (~2,017 tokens) and full ~24 rows (~9,480). What projecting the ~1,607-token code payload would cost is page yield — roughly 6 requirements per page instead of 24 — and therefore cost-per-answer:

tokens
a bounded hdf_query response~295
one requirement's code payload~1,607 (median; 929–2,566)
all 89 code payloads~149,735
the entire raw scan file~155,964

Returning every payload costs within 4% of handing the agent the raw file while still charging for the tool surface on top — so the read tools leave code unprojected and return normalized fields.

The boundary is specific to code, not to scanner data as a class. At verbosity: "full", hdf_query projects descriptions[] verbatim, and a converter may have placed equivalent scanner detail there — the Grype converter, for instance, puts related-vulnerability JSON in the check description. So a tool-specific question is sometimes answerable from this surface after all: the 45 grype requirements carrying related vulnerabilities can be counted from descriptions[label="check"] at full verbosity. Reach for full when you need that detail; fall back to the source file only for data no verbosity projects — the code payload itself, or a field the converter dropped.

The full-verbosity contract (decision x0bq): concise is the bounded default and excludes the payload; asking for full is opting into a larger, converter-determined payload of unknown size. The per-response token bound holds either way — full just fits fewer requirements per page. This documents existing behavior; it is not a change to it. A future revisit could add an explicit, bounded, per-requirement code fetch — never a widening of the default response.

Correlation fields ​

Normalization's quiet payoff is a shared vocabulary: findings from unrelated tools become correlatable once they carry the same normalized keys. Control-level correlation already works on the default surface — tags.nist is present on ~99% of requirements across sources, tags.cci on ~94% — and full verbosity returns both.

Below the control level, four normalized keys are the ones that meaningfully join findings across sources but that no read tool projects by default: cwe (weakness, bridges SAST and SCA), cvss (vulnerability scoring), affectedPackages (name/version/purl/fixedInVersion), and sourceLocation (file/line — the broadest cross-class key, spanning SAST, DAST, IaC, and secret findings). Each is class-scoped, not universal, so returning them by default would bloat every row for the findings that lack them.

hdf_query exposes them opt-in through a fields array, additive to whatever verbosity selects:

json
{ "source": { "path": "scan.json" },
  "fields": ["cwe", "cvss", "affectedPackages", "sourceLocation"] }

A requirement that lacks a requested field simply omits that key (never null), so a consumer joins on presence. An unknown field name is refused. The default (no fields) response is unchanged, and each key is added to a row only when the caller asks for it — so the correlation surface costs nothing until it is used.

Budget guidance for agent hosts ​

The HDF tools are analysis-heavy and naturally loopable (open → query → compliance → diff), so an agent host can run long, expensive tool-call sequences against this server. The server bounds the size of each individual response, but it cannot bound how many times a host calls it. The two budgets below are complementary: the server owns per-response sizing; the host owns loop control.

What the server enforces (per-response token ceilings) ​

Every response is a compact summary plus a handle — never a document body — and each collection-returning tool bounds its response against a serialized-size token budget (estimated at roughly four bytes per token). The ceilings are fixed constants, not configuration:

ToolResponse ceiling
hdf_open1,000 tokens
hdf_inspect2,000 (concise) / 10,000 (full)
hdf_query2,000 (concise) / 10,000 (full)
hdf_diff2,000 (concise) / 10,000 (full)
hdf_compliance2,000 tokens
hdf_aggregate2,000 tokens
hdf_validate2,000 tokens

concise is the default verbosity; request full only when you need raw content, and expect a larger response.

Separately, the one-time tools/list handshake at session start costs about 5,100 tokens for all ten tool schemas (a per-tool ceiling of 600 keeps any single schema in check). That is a fixed startup cost, paid once per server session, not per call — and you can trim it by advertising only the tools a deployment needs (see Advertising fewer tools).

A response is not paid once — it compounds. A tool result stays in the conversation and is re-sent on every subsequent turn, so a single full response (up to 10,000 tokens) in a five-turn exchange is charged roughly five times. For a long automated loop — the motivating case — that compounding dwarfs the tool surface. Prefer concise plus pagination over full when a run has many turns; reach for full only when a turn genuinely needs raw content. The server helps here in one specific way: each response carries its full payload once, in structuredContent, with only a short one-line gist in the content text block — so a host that feeds content to the model is not charged the payload twice.

Pagination and truncation ​

When a result set does not fit the ceiling, the tool does not fail and does not silently drop data. It returns a successful response with the rows that fit plus explicit bounding metadata:

json
{ "total": 412, "returned": 37, "truncated": true, "nextPage": 1,
  "notice": "returned 37 of 412; narrow with verbosity=concise or fetch page=N" }
  • total vs. returned tells you how much was withheld.
  • truncated: true with nextPage means more pages exist; pass page: N to fetch them.
  • notice always names the remedy — the narrowing parameter to use next (a filter, verbosity=concise, or page=N). Prefer narrowing with filters (hdf_query accepts status, severity, ID, tag, and free-text filters) over paging blindly through a large set.

What the host must enforce (loop ceilings) ​

The server is stateless with respect to an agent's loop: each request is handled independently, so it has no notion of "calls this turn" and cannot enforce a per-turn tool-call or token ceiling. That control belongs to the host wiring the agent onto the server. Recommended practice, mirroring the per-endpoint budget profiles MITRE's Caldera MCP plugin documents (max_tool_calls / max_tokens):

  • Cap tool calls per turn. The open → query → compliance → diff loop can iterate indefinitely; set a ceiling appropriate to the task so a runaway agent cannot spin unbounded.
  • Cap tokens per turn. Bound the cumulative response tokens the host will spend before it must stop and summarize, independent of the per-response ceilings above.
  • Treat these as host policy, not server guarantees. The server will keep answering well-formed calls; only the host can decide when enough is enough.

Workflow guidance ​

The generative workflows an agent runs against this server — drafting attestation amendments for notReviewed requirements, or applying a VEX document as amendments — ship as a usage Skill, not as MCP prompts. The repository ships an hdf-mcp-workflows skill carrying the tool sequencing each workflow needs, for use with coding agents that support skills.

Worked example — the HDF MCP by example ​

For a runnable, end-to-end primer on using this server well — wiring an MCP client over stdio, the normalize-once-then-query pipeline, read-tool sequencing, and the handle/source model — see the hdf-mcp-demo example-consumer project (a sibling to hdf-conmon-demo; it drives the shipped hdf binary, importing nothing private to this repo).

It also measures why the bounded, normalized surface is cheaper for an agent than reading raw scans into context. Across five real scan formats (SAST, DAST, vulnerability, SBOM inventory, and CCE compliance) it compares the tokens an HDF MCP response costs against the raw files — model-free and reproducible offline — and then puts the same questions to a real model, so the ideal saving can be checked against what an agent actually spends.

It reports the cases where HDF loses, including questions the bounded read surface cannot answer at all. Conversion is not lossy here: the scanner's original finding is preserved verbatim in the requirement's code. But no read tool projects code, so an agent asking for a tool-specific field must fall back to the raw scan and the normalized surface buys it nothing. See the repo for the current measured figures.

Released under the Apache 2.0 License.