Skip to content

Heimdall Data Format (HDF) v3.7.0 Specification ​

Version: 3.7.0 Schema: JSON Schema draft 2020-12 License: Apache-2.0 | The MITRE Corporation

Overview ​

HDF is a JSON format for representing security assessment data. It normalizes outputs from vulnerability scanners, compliance checkers, and configuration auditors into a unified data model. Seven document types cover the full assessment lifecycle, plus a continuous-monitoring change-event stream (section 8).

DocumentPurposeKey Fields
ResultsAssessment findings (pass/fail)baselines, components, statistics
BaselineRequirements sets (without findings attached)requirements, groups, inputs
SystemDescription of system under assessmentcomponents, dataFlows, controlDesignations
PlanAssessment planassessments, schedule
AmendmentsStatus overrides (waivers, POAMs)overrides
ComparisonDiff between assessmentsrequirementDiffs, summary
Evidence PackageAudit bundlecontents (references to other docs)

Documents reference each other via URI strings: systemRef, planRef.

Conventions ​

  • All documents use "unevaluatedProperties": false — unknown fields are rejected.
  • Date/time values use ISO 8601 (2026-01-15T10:30:00Z).
  • UUIDs use RFC 4122 format.
  • Checksums use {algorithm, value} pairs. Algorithms: sha256, sha384, sha512, blake3.
  • generator (optional on all documents): {name, version} of the producing tool.
  • labels (optional on most documents): {key: value} string map for flexible grouping. Well-known keys: system, component, environment, region, team.

Prose fields ​

Every HDF prose field — descriptions[].data, Requirement_Core.title and .code, result message and codeDesc, override reason, and every description, justification, explanation and summary — is a plain JSON string. Line breaks are \n inside that string. There is no array-of-lines form, and none will be added.

Prose is carried verbatim. Interior blank lines, CRLF, tabs, leading and trailing whitespace, and a trailing newline are all data. Nothing in HDF normalizes prose on ingest: a converter carries the source tool's text through unchanged, so prose survives a read-write round trip byte-for-byte. HDF normalizes structure, not words.

The promise is possible because HDF is JSON. XML (XML 1.0 §2.11) and YAML (YAML 1.2.2 §5.4) both require a parser to fold line breaks to LF, so no format in those families can make it; JSON mandates no line-ending normalization at all, so "a\r\nb" and "a\nb" are distinct strings that both round-trip unchanged. It is worth stating because edge whitespace is frequently the tool's own output framing — a real InSpec result message begins and ends with a newline — and because each exporter would otherwise invent its own trim rule. SARIF 2.1.0 §3.3.2 is the precedent for saying it out loud: a producer "SHALL preserve the text artifact's line breaking convention (for example, \n or \r\n)".

HDF declares no markup. Not per field, not per document; no markup or format property will be added. A prose value is whatever the source tool emitted. Real converter output in this repository carries plain text (InSpec, STIG), CommonMark (cyclonedx, neuvector, snyk, sarif) and raw HTML (msft-secure-score, aws-config, netsparker) in the same fields, and a merged multi-scanner document can hold all three inside one descriptions[] array. A consumer that renders HDF prose must therefore sanitize it and must not assume plain text; no HDF library renders prose itself.

Single-line labels versus prose. Short single-line labels are exactly descriptions[].label (declared at five schema sites: Requirement_Core.descriptions[], Requirement_Description, Evaluated_Requirement.descriptions[], and both Baseline_Requirement_Descriptions declarations), Milestone.title, Source.label, Per_Source_Summary.label, Annotation.label, and Scanner_Conflict.values[].sourceLabel. Everything else listed above is prose and carries no length, line-break or whitespace constraint. Requirement_Core.title is prose despite its name: it is a pass-through of a foreign tool's free-text field (InSpec control titles, XCCDF element text, Graph API titles), and real emitted values contain line breaks and trailing spaces. Baseline_Metadata.title and Requirement_Group.title are prose for the same reason. Milestone.title is the only field that currently carries a single-line pattern (minLength: 1 plus a pattern banning line breaks and edge whitespace), and that pattern is the precedent shape for any future label constraint.

Where normalization is legitimate. Only at an export boundary whose own type forbids what HDF allows, and only for that target. dev-docs/adr-0014-oscal-namespace-and-prose-carriage.md §1.7.1 collapses line terminators and trims edge whitespace for OSCAL StringDatatype sinks, carrying the exact HDF value in the accompanying remarks; that rule is scoped to those sinks and is not generalized. An OSCAL markup-multiline home must pass \n through untouched, because a blank line is the paragraph delimiter there. A target that folds line endings by specification (XCCDF, or any other XML sink) returns LF-normalized prose — the target format's documented loss, stated per exporter, not an HDF rule.

Evidence data encoding ​

Evidence carries supporting material on a requirement or override; its data field is required and must be non-empty (minLength: 1). The encoding field names how to read data — utf-8 for the raw text of a code/log entry, the URL string for a url entry, or base64 for a screenshot/file. When encoding is exactly base64 (lowercase), data must be bare base64 matching OSCAL's Base64Datatype (^[0-9A-Za-z+/]+={0,2}$): no data: URI prefix, no surrounding URL, and no line breaks. A file or screenshot is therefore the raw base64 of its bytes; a link to an external artifact is a url entry, not a base64 body. (data non-emptiness and the bare-base64 rule tightened in v3.7.0)

Cardinality invariants ​

Several arrays in the schema declare minItems: 1:

PathConstraint
Evaluated_Baseline.requirementsminItems: 1
Evaluated_Requirement.resultsminItems: 1
Evaluated_Requirement.descriptionsminItems: 1 (must include label default)
Baseline.requirementsminItems: 1
Baseline_Requirement.descriptionsminItems: 1 (must include label default)
Components (system)minItems: 1

Rationale. An HDF Results document records the outcome of an assessment. A baseline with zero requirements (or a requirement with zero results) is structurally meaningless — it implies the producer claims an assessment happened but recorded no work. The cardinality invariants force producers to commit to a non-empty record, and let consumers (Heimdall app, hdf query, dashboards, downstream OSCAL exporters) safely access requirements[0].results[0] and requirements[0].descriptions[0] without defensive null-checking. baselines itself carries no minItems: a results document may hold zero baselines, and a consumer indexing baselines[0] must check first. Empty arrays would also be semantically ambiguous: clean scan? filter excluded everything? truncated input? converter bug? An explicit record disambiguates. The constraint matches peer formats — OSCAL AssessmentResult and InSpec profiles have equivalent non-empty-collection guarantees.

Migration note. Legacy HDF (the InSpec-ExecJSON-shaped output that the original Heimdall2 mappers produced) did not enforce these cardinality invariants. A clean Heimdall2 scan emitted profiles[0].controls: [] and the consumer was expected to tolerate empty arrays. The v3 schema deliberately tightened the contract: producers MUST emit non-empty arrays, and clean scans MUST synthesize a passed placeholder per the convention below. Consumers reading v3 HDF can rely on first-element access being safe; producers translating from legacy HDF v2 (InSpec-shaped) data MUST add synthesis at the converter boundary.

Clean-scan convention: synthesize a passed placeholder ​

A direct consequence of the cardinality invariants: a scanner that runs cleanly (zero findings) still must produce at least one requirement and one result. Converters MUST synthesize a placeholder rather than emitting an empty array.

Shape:

FieldValue
id<source>-no-findings (kebab-case converter source name) — or <sub-tool>-no-findings for multi-baseline converters where each baseline represents a distinct sub-tool (e.g. SARIF per-run, MSDO per-scanner)
title"No findings reported"
impact0
tags{}
descriptions[0]{label: "default", data: <codeDesc>}
results[0].statuspassed
results[0].codeDesc<Tool> <verb> <target> and reported zero <noun>. (verb: scanned default, analyzed for SCA tools, ran for multi-tool wrappers; noun: findings default, vulnerable components for SCA/dependency/container-vuln tools)
results[0].startTimescan timestamp from source, or current time if unavailable

Spec-backed semantics for passed. This is not a "no information" placeholder — it is a positive assertion that the scan ran against in-scope inputs and detected nothing wrong. The framing aligns with: NIST SP 800-53A Satisfied, OSCAL satisfied, XCCDF pass, STIG Not_a_Finding, and SARIF v2.1.0 §3.7.2 ("an empty results array shall be interpreted to mean that the analysis tool did not detect any results").

Distinguish from notApplicable. Synthesize notApplicable (not passed) when the rule's applicability check itself didn't fire — e.g. a cloud-config rule with zero in-scope resources, where the rule never evaluated against anything. The two convey materially different audit positions: passed says "we checked and you're compliant"; notApplicable says "we didn't check this one." Producers SHOULD use whichever the source format's semantics support; consumers SHOULD treat them as distinct.

Reference helper implementations: shared.BuildNoFindingsRequirement (Go) and buildNoFindingsRequirement (TypeScript) in hdf-converters/shared/. Every in-tree converter that emits a passed placeholder for clean scans calls one of these.


1. Results ​

Assessment findings from running security checks against target systems.

Schema ID: https://mitre.github.io/hdf-libs/schemas/hdf-results/v3.7.0

Top-Level Fields ​

A Results document is the primary output of HDF converters. It captures what was scanned (components), what was checked (baselines with requirements), and what happened (results with pass/fail status). The tool field identifies the original security tool; generator identifies the converter that produced the HDF file. Cross-document links (systemRef, planRef) connect results to the broader assessment context.

FieldTypeRequiredDescription
baselinesEvaluated_Baseline[]yesBaselines evaluated with findings
componentsComponent[]noSystem components assessed
statisticsStatisticsnoDuration, result counts
timestampdate-timenoWhen assessment ran
toolToolnoSecurity tool that produced scan data ({name, version, format})
generatorGeneratornoTool that produced this HDF file
runnerRunnernoExecution environment (distinct from components)
integrityIntegritynoCryptographic integrity metadata
preAmendmentChecksumChecksumnoHash of this document as it stood immediately before the most recent amendment application, written by hdf amend apply. Covers the results file's raw bytes as read, not a canonical form. Records only the most recent step: re-applying replaces it rather than accumulating. Distinct from the override-level previousChecksum, which chains amendments to each other over a canonical form (v3.6.0)
systemRefURI-referencenoLink to System document
planRefURI-referencenoLink to Plan document
extensionsobjectnoTool-specific metadata
idUUIDnoUnique assessment run identifier
remediationRemediationnoReference to automated fix resources
externalReferencesExternal_Reference[]noInert references to external artifacts (CTI/STIX, BOMs, advisories, runbooks) relevant to the assessment as a whole (v3.5.0)
derivationDerivationnoLineage of a reconciled result set reassembled from a seed snapshot plus change events (see Section 8); absent on documents produced directly by a scan (v3.5.0)

Evaluated_Baseline ​

An evaluated baseline represents the results of a set of security tests that have been run against a target system, aligned back to a baseline. It contains the original baseline metadata (name, version, author information) plus the assessed requirements with their pass/fail results. For example, an Evaluated Baseline could represent a Security Content Automation Protocol (SCAP) benchmark after it is evaluated. In InSpec terms, an Evaluated Baseline is a profile after it has been executed against a target with inspec exec and produced results. HDF converters should produce one evaluated baseline per scan profile, scanner module, or logical grouping of checks in the source tool's output.

FieldTypeRequiredDescription
namestringyesUnique baseline name
requirementsEvaluated_Requirement[]yesAssessment findings; minItems: 1
titlestringnoHuman-readable title
versionstringnoBaseline version
descriptionstringnoDetailed description
maintainerstringnoMaintainer name/contact
summarystringnoe.g. STIG header
copyrightstringnoCopyright holder(s)
copyrightEmailstringnoCopyright contact email
licensestringnoe.g. "Apache-2.0"
statusstringnoe.g. "loaded"
statusMessagestringnoExplanation of status
groupsRequirementGroup[]noLogical groupings of requirements
inputsInput[]noTyped parameters for execution
dependsDependency[]noBaseline dependencies
supportsSupportedPlatform[]noSupported platform targets
integrityIntegritynoCryptographic integrity metadata for this baseline
originalChecksumChecksumnoImmutable baseline definition hash
resultsChecksumChecksumnoRaw results hash (before amendments)
parentBaselinestringnoParent baseline name (overlay/wrapper)
labelsnoKey-value grouping metadata
externalReferencesExternal_Reference[]noInert references to external artifacts relevant to the baseline definition (CTI/STIX, advisories, source catalogs) (v3.5.0)
extensionsobjectnoTool-specific metadata

Evaluated_Requirement ​

A single security requirement with test results. Each requirement maps to one testable security control — an InSpec control, a SonarQube rule, a Nessus plugin, or an OSCAL control objective. The impact score (0.0–1.0) indicates severity, tags carry framework mappings (NIST SP 800-53, CCI, CWE), and results holds the individual test executions. Multiple results occur when the same check runs against multiple resources or when findings are grouped by rule.

FieldTypeRequiredDescription
idstringyesRequirement identifier, e.g. "SV-238196"
impactnumberyesSeverity score 0.0–1.0
tagsyesMetadata; common keys: nist, cci, severity, cwe
resultsRequirementResult[]yesTest executions; minItems: 1
descriptionsDescription[]yesMust include label "default"; minItems: 1
titlestringnoHuman-readable title
codestringnoSource code of the check
severitySeveritynoExplicit severity rating
effectiveStatusResultStatusnoStatus after applying overrides
sourceLocationSourceLocationnoFile path and line number
refsReference[]noExternal document references
statusOverridesStatusOverride[]noAudit trail of status changes
poamsPoamElement[]noRemediation tracking (don't change status)
evidenceEvidence[]noScreenshots, logs, code samples
dispositionOverrideTypenoOverride disposition applied to this requirement
effectiveImpactnumbernoEffective impact after overrides (0.0–1.0)
effectiveChecksumChecksumnosha256 over the resolved effective posture ({status, impact, disposition}) for per-control change detection; flips only when the operative posture changes (v3.5.0)
controlTypeControlTypenoNIST SP 800-53 / SP 800-53A categorization: policy | procedure | technical | management | operational (v3.2.0)
verificationMethodVerificationMethodnoHow this requirement is verified: automated | manual-by-design | manual-pending-automation | hybrid. Disambiguates the two cases that null code previously overloaded (v3.2.0)
applicabilityApplicabilitynoWithin-baseline applicability: required | optional | advisory. Distinct from severity (risk weight) and status (lifecycle) (v3.2.0)
cvssCVSS[]noTyped CVSS scoring for the finding. Multi-entry to handle multi-CVE findings; all four major CVSS versions supported (v3.3.0)
epssEPSSnoEPSS exploit-probability data (percentile + score) (v3.3.0)
kevKevnoCISA Known Exploited Vulnerabilities catalog status (v3.3.0)
cwestring[]noCWE classification IDs (e.g. CWE-79). Replaces free-form tags.cwe (v3.3.0)
affectedPackagesAffected_Package[]noAffected-package identifiers (ecosystem + name + version) for vulnerability findings (v3.3.0)
externalReferencesExternal_Reference[]noInert references to external artifacts relevant to this requirement (CTI/STIX correlation, advisories, control/definition sources) (v3.5.0)

RequirementResult ​

A single test execution within a requirement. Each result records what was tested (codeDesc), when (startTime), how long it took (runTime), and the outcome (status). For failed tests, message explains the failure. For errors, exception and backtrace capture the fault. The resource and resourceId fields identify the specific system resource that was tested (e.g. a file, service, or package).

FieldTypeRequiredDescription
statusResultStatusyesTest outcome
codeDescstringyesDescription of what was tested
startTimedate-timeyesWhen test started
runTimenumbernoExecution time in seconds
messagestringnoExplanation of result
exceptionstringnoException type if error occurred
resourcestringnoResource type tested, e.g. "file", "service"
resourceIdstringnoResource identifier, e.g. "/etc/passwd"
backtracestring[]noStack trace if exception occurred

Component ​

The system element that was assessed. Components are polymorphic — each has a type discriminator that determines which additional fields are available. All components share name, type, and optional fields for identity (componentId), ownership (owner), external cross-references (externalIds), labels, BOM attachment (boms[] — SBOM, AI-model, or dataset), artifact integrity (integrity[]), baseline references, per-system input overrides, and a target selector. Components in Results are typically populated with minimal fields by converters; components in System documents carry the full set including componentId for cross-document correlation.

FieldTypeRequiredDescription
namestringyesHuman-readable component name
typeComponentTypeyesDiscriminator (see types below)
componentIdUUIDnoStable identity for cross-document correlation
descriptionstringnoComponent role or purpose
ownerIdentitynoTeam or individual responsible for this component
externalIdsnoExternal ID map (aws, azure, cmdb, emass)
labelsnoKey-value grouping metadata
bomsBillOfMaterials[]noComponent-scoped BOMs (SBOM, ai-model, dataset, or reserved bomType), by passthrough or normalized. Replaces the former sbom/sbomRef/sbomFormat trio (v3.4.0)
integrityChecksum[]noCryptographic integrity of the component's artifact (model weights/shards, dataset archive, image, package bytes). Generic home replacing per-type digest/checksum (v3.4.0)
baselineRefsstring[]noNames of baselines that apply
inputOverridesInputOverride[]noSystem-specific overrides for baseline input values
targetSelectorTargetSelectornoLabel selector matching targets that belong to this component

Type-specific fields:

TypeExtra Fields
hosthostname, fqdn, domain, ipAddress, macAddress, osName, osVersion
containerImageimageId, registry, repository, tag
containerInstancecontainerId, image, runtime
containerPlatformplatformType, clusterName, namespace, version
cloudAccountprovider, accountId, region
cloudResourceprovider, resourceType, resourceId, arn, region
repositoryurl, branch, commit
applicationurl, version, environment
artifactpackageManager, packageName, version
networkcidr, gateway
databaseengine, version, host, port
aiModel (v3.4.0)modelId, version (detail in the attached ai-model BOM)
dataset (v3.4.0)datasetId, version (detail in the attached dataset BOM)

2. Baseline ​

Security requirements without results (before assessment).

Schema ID: https://mitre.github.io/hdf-libs/schemas/hdf-baseline/v3.7.0

Shares the baseline metadata fields with Evaluated_Baseline but uses Baseline_Requirement (no results, no effectiveStatus) instead of Evaluated_Requirement. The evaluation-only fields of Evaluated_Baseline (description, statusMessage, parentBaseline, originalChecksum, resultsChecksum, extensions) are not defined on a Baseline document.

FieldTypeRequiredDescription
namestringyesUnique baseline name
requirementsBaseline_Requirement[]yesSecurity requirements; minItems: 1
groupsRequirementGroup[]noLogical groupings of requirements
inputsInput[]noTyped parameters for execution
dependsDependency[]noBaseline dependencies
integrityIntegritynoCryptographic integrity metadata for this baseline
remediationRemediationnoReference to automated fix resources
generatorGeneratornoTool that produced this file
titlestringnoHuman-readable title
versionstringnoBaseline version
maintainerstringnoMaintainer name/contact
summarystringnoe.g. STIG header
copyrightstringnoCopyright holder(s)
copyrightEmailstringnoCopyright contact email
licensestringnoe.g. "Apache-2.0"
statusstringnoe.g. "loaded"
supportsSupportedPlatform[]noSupported platform targets
labelsnoKey-value grouping metadata
externalReferencesExternal_Reference[]noInert references to external artifacts relevant to the baseline definition (CTI/STIX, advisories, source catalogs) (v3.5.0)

Baseline_Requirement ​

A security requirement before assessment. Shares Requirement_Core with Evaluated_Requirement but carries none of the evaluation-side fields: no results, effectiveStatus, effectiveImpact, effectiveChecksum, disposition, statusOverrides, poams, or evidence, and none of the vulnerability-finding fields (cvss, epss, kev, cwe, affectedPackages), which exist only on Evaluated_Requirement. Used to represent standalone baselines (STIG profiles, CIS benchmarks) and in OSCAL catalog/profile conversions where no assessment has occurred yet.

FieldTypeRequiredDescription
idstringyesRequirement identifier
impactnumberyesSeverity score 0.0–1.0
tagsyesMetadata; common keys: nist, cci, severity
descriptionsDescription[]yesMust include label "default"; minItems: 1
titlestringnoHuman-readable title
codestringnoSource code of the check
severitySeveritynoExplicit severity rating
sourceLocationSourceLocationnoFile path and line number
refsReference[]noExternal document references
controlTypeControlTypenoNIST SP 800-53 / SP 800-53A categorization: policy | procedure | technical | management | operational (v3.2.0)
verificationMethodVerificationMethodnoHow this requirement is verified: automated | manual-by-design | manual-pending-automation | hybrid (v3.2.0)
applicabilityApplicabilitynoWithin-baseline applicability: required | optional | advisory (v3.2.0)
externalReferencesExternal_Reference[]noInert references to external artifacts relevant to this requirement (CTI/STIX correlation, advisories, control/definition sources) (v3.5.0)

3. System ​

Describes a system under assessment. A system document defines the authorization boundary, including what components make up the system, its security categorization (FIPS 199), and its authorization status (ATO). This corresponds to a FedRAMP system or an OSCAL SSP's system characteristics. Results and amendments reference the system via systemRef to establish which system the assessment applies to.

Schema ID: https://mitre.github.io/hdf-libs/schemas/hdf-system/v3.7.0

FieldTypeRequiredDescription
namestringyesSystem name
componentsComponent[]yesSystem components
systemIdUUIDnoStable identity for cross-document correlation independent of file location; expected in production documents
ownerIdentitynoTeam or individual responsible for the system's authorization and compliance (OSCAL system-owner responsible-party)
identifierstringnoSystem identifier (e.g. FedRAMP ID)
identifierSchemestringnoIdentifier scheme (e.g. "fedramp")
descriptionstringnoSystem description
authorizationStatusAuthorizationStatusnoATO status
authorizationDatedatenoDate of authorization decision
categorizationLevelCategorizationLevelnoFIPS 199 categorization
boundaryDescriptionstringnoSystem boundary narrative
controlDesignationsControlDesignation[]noControl inheritance declarations
dataFlowsDataFlow[]noData flows between components or external endpoints
bomsBillOfMaterials[]noSystem-scoped BOMs whose subject is the authorization boundary rather than one component (e.g. a SaaSBOM of services, a KBOM of cluster inventory, an OBOM); component-scoped BOMs attach on the component (v3.4.0)
labelsnoKey-value grouping metadata
integrityIntegritynoCryptographic integrity metadata for this document
versionstringnoDocument version
generatorGeneratornoTool that produced this file
externalReferencesExternal_Reference[]noInert references to external artifacts describing the system's threat environment or context (CTI/STIX, BOMs, advisories) (v3.5.0)

System Component ​

System components use the same polymorphic Component type as Results (see Section 1), with componentId required for stable cross-document references, data flow endpoints, and control designation binding.

Data Flow ​

Describes a data flow between components within the system or to external endpoints. Used for authorization boundary diagrams and data flow documentation. The to endpoint is one of: a local componentId (UUID), a Cross_System_Reference ({systemRef, componentId} naming a component in another system document), or an External_Endpoint ({external: true, description} for an endpoint outside all modeled systems).

FieldTypeRequiredDescription
fromUUIDyesSource componentId
toUUID | CrossSystemReference | ExternalEndpointyesDestination (local component, component in another system, or external endpoint)
protocolstringnoProtocol (e.g. "https", "grpc", "jdbc")
portintegernoPort number (1–65535)
directionDirectionnounidirectional (from→to only) or bidirectional (e.g. request/response)
descriptionstringnoFlow description
authenticationstringnoAuthentication mechanism for the connection (e.g. "mTLS", "OAuth2", "Kerberos")

Control Designation ​

Declares a control's designation within the system — whether it is common (provided by another component or external system), system-specific (implemented locally), or hybrid (shared responsibility). This maps to NIST SP 800-53 Appendix C control designations and OSCAL SSP by-component provided/inherited semantics. When providedBy is set, the control is provided by a local component; when systemRef is set, the provider is in another system's authorization boundary. When neither is set, the provider is external with no formal HDF representation (e.g. a cloud provider's physical security).

FieldTypeRequiredDescription
controlIdstringyesControl identifier (e.g. "SC-7", "AC-2 (1)")
designationControlDesignationTypeyescommon, system-specific, or hybrid
descriptionstringyesJustification for this designation
providedByUUIDnocomponentId of local provider
systemRefURI-referencenoReference to external system document
inheritedByUUID[]nocomponentIds that inherit (default: all)

4. Plan ​

Assessment plan defining what to assess and how. A plan document describes the scope, methodology, and schedule for an upcoming security assessment. It references the system under test via systemRef and lists the individual assessments to be performed. This corresponds to an OSCAL SAP (Security Assessment Plan) or a FedRAMP test plan.

Schema ID: https://mitre.github.io/hdf-libs/schemas/hdf-plan/v3.7.0

FieldTypeRequiredDescription
namestringyesPlan name
assessmentsAssessment[]yesPlanned assessment activities
planIdUUIDnoStable identity for this plan; expected in production documents
typePlanTypenoAssessment methodology
descriptionstringnoPlan description
systemRefURI-referencenoLink to System document
scheduleSchedulenoAssessment schedule
labelsnoKey-value grouping metadata
integrityIntegritynoCryptographic integrity metadata for this document
versionstringnoDocument version
generatorGeneratornoTool that produced this file
externalReferencesExternal_Reference[]noInert references to external artifacts relevant to the plan (CTI/STIX, advisories, methodology documents) (v3.5.0)

Assessment ​

A single assessment within a plan — defines which baseline to run against which component.

FieldTypeRequiredDescription
baselineRefURI-referenceyesReference to baseline to evaluate
componentRefUUIDnocomponentId of target component (direct binding)
targetSelectorTargetSelectornoLabel selector for target matching
inputsobjectnoResolved input values
runnerRunnerConfignoRunner/scanner configuration
descriptionstringnoAssessment purpose

5. Amendments ​

Status overrides applied after assessment (waivers, attestations, POAMs).

Schema ID: https://mitre.github.io/hdf-libs/schemas/hdf-amendments/v3.7.0

FieldTypeRequiredDescription
namestringyesAmendment set name
overridesOverride[]yesStatus overrides
amendmentIdUUIDnoStable identity for this amendments document; distinguishes several amendment documents targeting the same results
descriptionstringnoDescription of this amendment set
systemRefURI-referencenoLink to System document
appliedByIdentitynoWho applied the overrides
approvedByIdentitynoWho approved the overrides
labelsnoKey-value grouping metadata
integrityIntegritynoCryptographic integrity metadata for this document
signatureSignaturenoDigital signature
versionstringnoDocument version
generatorGeneratornoTool that produced this file

Override ​

A deliberate change to an assessed requirement's compliance status. Waivers grant temporary acceptance of a known risk. Attestations assert manual verification of a requirement that cannot be automatically tested. POAMs track planned remediation with milestones. False positives mark findings confirmed as not applicable. Risk adjustments modify severity based on contextual analysis. Operational requirements document accepted deviations due to mission needs. Inherited overrides indicate that a control is provided by another component or system (typically overriding to notApplicable or passed). Each override records who authorized it, why, and when it expires. The previousChecksum field chains each override to the one before it, making an override edited in place after the fact detectable by hdf amend verify.

FieldTypeRequiredDescription
typeOverrideTypeyesOverride category
requirementIdstringyesID of requirement being overridden
reasonstringyesFree-text auditor-readable rationale for the override
appliedByIdentityyesIdentity of who applied this amendment
appliedAtdate-timeyesWhen this amendment was applied
expiresAtdate-timeyesWhen this amendment expires. No permanent amendments
statusResultStatusconditionalNew effective status. At least one of status or impact must be set, except when type: operationalRequirement (which forbids both)
impactImpact_OverrideconditionalOverride to the requirement's impact score ({value: 0.0–1.0}). At least one of status or impact must be set; forbidden when type: operationalRequirement
baselineRefstringnoName of the baseline containing the requirement; needed when the system has multiple baselines with overlapping IDs
componentRefUUIDnocomponentId this amendment is scoped to; omit for system-wide
inheritedFromUUIDnocomponentId of the local control provider; primarily used with type: inherited
affectedPackagesAffected_Package[]noSoftware packages this amendment is scoped to (purl/cpe/name+version), distinct from componentRef
evidenceEvidence[]noSupporting evidence (screenshots, logs, URLs, documents)
signatureSignaturenoDigital signature for non-repudiation
previousChecksumChecksumnoChecksum of the prior amendment; detects an in-place edit of any amendment that has a later one chained to it
cvssCVSSnoStructured CVSS scoring backing this override; on riskAdjustment, impact.value should be approximately cvss.computedScore / 10.0 (v3.3.0)
justificationJustificationnoStructured controlled-vocabulary classification (VEX-aligned; see the Justification enumeration). Complements (does not replace) reason (v3.3.0)
milestonesMilestone[]noRemediation milestones (primarily for poam type)
externalReferencesExternal_Reference[]noInert references to the external artifacts behind this override (e.g. the STIX bundle/object motivating an E:H riskAdjustment); distinct from evidence (v3.5.0)

6. Comparison ​

Diff between two or more assessment documents. A comparison captures how compliance posture changed between scans, across environments, or between baseline versions. The comparisonMode indicates the type of analysis (temporal drift, fleet comparison, baseline evolution, etc.). Each requirement diff records whether a control is new, absent, fixed, regressed, or unchanged. Comparisons are produced by hdf diff and consumed by dashboards to show trend data.

Schema ID: https://mitre.github.io/hdf-libs/schemas/hdf-comparison/v3.7.0

FieldTypeRequiredDescription
formatVersionstringyesComparison format version
comparisonModeComparisonModeyesType of comparison
sourcesSource[]yesDocuments being compared
summarySummaryyesAggregate diff statistics
requirementDiffsRequirementDiff[]yesPer-requirement diffs
timestampdate-timenoWhen comparison was generated
generatorGeneratornoTool that produced this file
matchingMatchingConfignoHow requirements were matched
baselineDiffsBaselineDiff[]noPer-baseline diffs
componentDiffsComponentDiff[]noPer-component diffs
packageDiffsPackageDiff[]noPer-package diffs
systemRefURI-referencenoLink to System document
driftDriftAnalysisnoDrift analysis metadata
annotationsAnnotation[]noHuman/tool annotations on diffs
integrityIntegritynoCryptographic integrity metadata
extensionsobjectnoTool-specific metadata
externalReferencesExternal_Reference[]noInert references to external artifacts relevant to the comparison (CTI/STIX, advisories) (v3.5.0)

7. Evidence Package ​

Bundles references to assessment artifacts for audit and compliance submission. An evidence package collects results, baselines, amendments, system descriptions, and supporting materials (screenshots, logs, SBOMs) into a single auditable unit. This corresponds to a FedRAMP security package or an OSCAL POA&M submission bundle. The completenessCheck field validates that all expected artifacts are present.

Schema ID: https://mitre.github.io/hdf-libs/schemas/hdf-evidence-package/v3.7.0

FieldTypeRequiredDescription
namestringyesPackage name
contentsContentReference[]yesReferences to bundled artifacts; minItems: 1
packageIdUUIDnoStable identity for this package; expected in production ATO submissions
descriptionstringnoPackage description
systemRefURI-referencenoLink to System document
planRefURI-referencenoLink to the Plan document that drove the assessment; used for completeness verification
preparedByIdentitynoWho prepared this package
preparedAtdate-timenoWhen package was prepared
externalEvidenceExternalEvidenceReference[]noReferences to external native-format evidence (log/telemetry corpora) by URI + hash + format, never transcoded into HDF (v3.4.0)
completenessCheckCompletenessChecknoValidation of package contents
signatureSignaturenoDigital signature
labelsnoKey-value grouping metadata
integrityIntegritynoCryptographic integrity metadata for this document
versionstringnoDocument version
generatorGeneratornoTool that produced this file
externalReferencesExternal_Reference[]noInert references to external artifacts relevant to the package beyond the document references it already carries (v3.5.0)

Content Reference ​

A reference to an HDF document or BOM/manifest document included in the evidence package.

FieldTypeRequiredDescription
typeContentTypeyesDocument type; bom covers any Bill-of-Materials/manifest document, whose specific kind is carried by the referenced document's bomType
uriURI-referenceyesDocument location
checksumChecksumnoDocument integrity hash
descriptionstringnoEntry description
componentRefUUIDnocomponentId this content relates to

External Evidence Reference (v3.4.0) ​

A reference to external native-format evidence (a log or telemetry corpus, or another artifact) carried by URI, integrity hash, and a format discriminator. Reference-only: the artifact is never embedded (corpora can be huge) or transcoded (that would be lossy); it stays canonical in its native format and HDF acts as the structured index. Reserved format values are open, serialized standards; custom or vendor-specific kinds use an x- prefix (e.g. x-splunk-export) so they cannot collide with a value later promoted into the reserved set.

FieldTypeRequiredDescription
uriURI-referenceyesLocation of the native-format artifact
formatExternalEvidenceFormatyesNative format discriminator: ecs | ocsf | cyclonedx | spdx | raw-log, or an x- prefixed custom value
checksumChecksumnoIntegrity hash of the referenced artifact
mediaTypestringnoIANA media type of the on-disk serialization (e.g. application/x-ndjson), orthogonal to format
formatVersionstringnoProducer-declared format version (e.g. ECS 9.4.0); free text, not validated against a registry
descriptionstringnoEntry description
metadataExternalEvidenceMetadatanoDescriptive metadata: recordCount (integer), timeRange ({start, end} date-times), collector (string). Does not affect integrity

8. Requirement Change Event (v3.5.0) ​

A single continuous-monitoring event describing how one requirement's posture changed between two same-target results scans. Producers (hdf events derive) emit an NDJSON stream of these events; a batch replays onto a seed results document to reassemble a reconciled state (hdf events apply) or folds into a systemDrift comparison (hdf events fold). The unit is the requirement, keyed by (systemRef, componentId, requirementId) with a per-key integer sequence as the sole ordering authority. The SIEM export guide documents the change-event projections to OCSF and SARIF.

Schema ID: https://mitre.github.io/hdf-libs/schemas/hdf-requirement-change-event/v3.7.0

Envelope (CloudEvents-grounded dedup + ordering):

FieldTypeRequiredDescription
eventIdUUIDyesUUIDv5 over the fixed namespace + entity key + sequence; dedup identity
sourceURIyesProducer URI recorded on every event
sequenceinteger (≥0)yesPer-key monotonic order — the sole ordering authority
systemRefstringyesSystem reference component of the entity key
componentIdUUIDyesComponent component of the entity key
timestampdate-timeyesEvent occurrence (the next document's scan time)
priorChecksumChecksum | nullyeseffectiveChecksum of the prior state; chains events for a key
schemaRefURInoOptional schema reference stamped on the event

Event body:

FieldTypeRequiredDescription
requirementIdstringyesThe requirement whose posture changed
stateEvent_Requirement_Stateyesnew | absent | updated | fixed | regressed
beforeEffective_Projection | nullyesThin prior posture (effectiveStatus, effectiveImpact); null for new
afterEvaluated_RequirementyesFull after-state (absent except for a tombstone absent event)
changeReasonsChange_Reason[]noWhy the posture changed (same vocabulary as the comparison diff)

Derivation (v3.5.0) ​

Lineage stamped on a reconciled Results document (one produced by hdf events apply rather than by a scanner) so it can never masquerade as primary scan evidence. Records exactly which seed snapshot and event horizon the document represents; conceptually a PROV qualified derivation. Carried as derivation on the Results root; documents produced directly by a scan omit it.

FieldTypeRequiredDescription
seedobjectyesThe seed snapshot: {uri, checksum} (URI-reference plus Checksum of the seed document)
sourceURI-referenceyesEvent-stream producer context whose events were applied (matches the events' envelope source)
throughSequenceinteger (≥0)yesEvent watermark: the highest per-key sequence applied. A full-scan document supersedes the reconciled view as of its scan time; between scans, the reconciled document with the highest watermark is current
eventsAppliedinteger (≥0)yesNumber of change events applied to the seed
asOfdate-timeyesPosture-as-of time of the reconciled view (the analog of PROV generatedAtTime)

Enumerations ​

ResultStatus ​

passed | failed | notApplicable | notReviewed | error

Severity ​

critical | high | medium | low | informational

Impact Mapping ​

Impact is a float 0.0 to 1.0. Conventional mapping to severity:

ImpactSeverity
0.9critical
0.7high
0.5medium
0.3low
0.0informational / notApplicable

OverrideType ​

waiver | attestation | poam | inherited | falsePositive | riskAdjustment | operationalRequirement

PlanType ​

automated | manual | hybrid

ComparisonMode ​

temporal | baseline | fleet | multiSource | baselineEvolution | systemDrift

RequirementState (comparison) ​

new | absent | unchanged | updated | fixed | regressed | moved | split | merged

AuthorizationStatus ​

authorized | denied | pendingAuthorization | conditionallyAuthorized | notYetRequested | revoked

CategorizationLevel ​

low | moderate | high

ComponentType ​

host | containerImage | containerInstance | containerPlatform | cloudAccount | cloudResource | repository | application | artifact | network | database | aiModel | dataset

Direction (data flow) ​

unidirectional | bidirectional

ControlDesignationType ​

common | system-specific | hybrid

HashAlgorithm ​

sha256 | sha384 | sha512 | blake3

Justification (override) ​

VEX-aligned controlled vocabulary. The first five values are common to OpenVEX, CSAF VEX, and CycloneDX VEX; the remaining six are CycloneDX-specific and describe why the vulnerable code path is unreachable in the deployed configuration. The enum is extended additively across schema versions, so consumers should validate against the schema version the document declares.

component_not_present | vulnerable_code_not_present | vulnerable_code_not_in_execute_path | vulnerable_code_cannot_be_controlled_by_adversary | inline_mitigations_already_exist | requires_configuration | requires_dependency | requires_environment | protected_by_compiler | protected_at_runtime | protected_at_perimeter

ContentType (evidence package) ​

hdf-system | hdf-baseline | hdf-plan | hdf-results | hdf-amendments | hdf-comparison | bom

ExternalEvidenceFormat (evidence package) ​

ecs | ocsf | cyclonedx | spdx | raw-log | x-<custom>

BomType ​

Reserved, CycloneDX-aligned set: sbom | ai-model | dataset | hbom | cbom | saasbom | obom | mbom | kbom, or an x- prefixed custom value. sbom, ai-model, and dataset have normalized extensions; the rest are carried by passthrough only.


Shared Primitives ​

Checksum ​

json
{ "algorithm": "sha256", "value": "abc123..." }

Generator ​

json
{ "name": "sarif-to-hdf", "version": "1.0.0" }

Tool ​

json
{ "name": "Semgrep", "version": "1.34.0", "format": "SARIF" }

Identity ​

json
{ "type": "email", "identifier": "compliance@agency.gov" }

Types: email, username, system, agent, simple, other. Use email for email addresses, username for user accounts, system for deterministic non-interactive automation (CI jobs, cron, scanners), agent for an AI/LLM agent acting with autonomy (kept distinct from system so AI-generated provenance can be audited separately), simple for a basic string identifier, or other for a custom identity system.

Description ​

json
{ "label": "default", "data": "The system must..." }

Labels: default (required), check, fix, rationale (conventional). label is a short single-line label; data is prose carried verbatim with no declared markup — see Prose fields.

RequirementGroup ​

json
{ "id": "controls/ssh.rb", "title": "SSH Controls", "requirements": ["SV-1", "SV-2"] }

Dependency ​

json
{ "name": "os-baseline", "url": "https://...", "status": "loaded" }

Input ​

Typed parameters for baseline execution.

json
{
  "name": "max_sessions",
  "description": "Maximum concurrent sessions",
  "type": "Numeric",
  "value": 10,
  "constraints": { "operator": "le", "limit": 20 }
}

Input types: String | Numeric | Boolean | Array | Hash | Regexp. Comparison operators: eq | ne | lt | le | gt | ge | contains | matches | in | notIn.

SourceLocation ​

json
{ "ref": "controls/ssh.rb", "line": 42 }

SupportedPlatform ​

json
{ "platformName": "ubuntu", "platformFamily": "debian", "release": "22.04" }

Integrity ​

Document-level tamper evidence, available on every document root (and on evaluated baselines). algorithm and checksum are dependent: if one is present the other is required.

json
{ "algorithm": "sha256", "checksum": "abc123...", "signature": "...", "signedBy": "compliance@agency.gov" }

InputOverride ​

System-specific override of a baseline input value, carried on a component's inputOverrides[]. inputName and value are required; baselineRef scopes the override to one baseline (omit to apply to every baseline defining that input).

json
{ "baselineRef": "rhel9-stig", "inputName": "max_sessions", "value": 5, "justification": "Kiosk hosts", "approvedBy": { "type": "email", "identifier": "issm@agency.gov" } }

TargetSelector ​

Label selector matching targets by label key-value pairs; all listed labels must match (AND). Used on components (targetSelector) and plan assessments.

json
{ "labels.component": "WebTier" }

BillOfMaterials (v3.4.0) ​

One extensible Bom shape, discriminated by bomType, carries any manifest kind either by passthrough (ref or document, the native manifest untouched) or normalized into a queryable extension. Subject identity (name, version, componentId) is inherited from the host component or system; the BOM never re-owns it. A BOM must carry at least one of ref, document, packages, model, or dataset, and each normalized extension is permitted only on its matching bomType. Attached via boms[] on components and on the System root.

FieldTypeRequiredDescription
bomTypeBomTypeyesManifest kind (see the BomType enumeration)
formatstringyesSource manifest format, free-form (e.g. cyclonedx, cyclonedx-ml, spdx, spdx-ai, huggingface, croissant)
refURI-referencenoPassthrough by reference to the native manifest document
documentobjectnoPassthrough by embedding: the native manifest carried opaquely
hashesChecksum[]noIntegrity of the BOM document itself (not of its subject artifact, which is the component's integrity[])
uniqueIdstringnoStable identifier of the BOM document (CycloneDX serialNumber, SPDX documentNamespace)
licensestring | nullnoLicense of the BOM document as a whole (SPDX expression)
packagesSBOMPackage[]noNormalized sbom extension: flattened package inventory; each entry {name (required), version, purl, licenses[]}
modelAIModelExtensionnoNormalized ai-model extension: modelArchitecture, parameterCount, serializationFormat, adaptationType (finetune | adapter | quantized | merge), baseModelRef, datasetRefs[], intendedUse, learningApproach, task, performanceMetrics[], hyperparameters[], inputOutput
datasetDatasetExtensionnoNormalized dataset extension: recordCount, datasetFormat, dataClassification, intendedUse, modality, provenance, statisticalProperties, baseDatasetRefs[], derivation (filtered | augmented | merged | sampled)

External_Reference (v3.5.0) ​

A generalized, purpose-agnostic reference to any external artifact by identity and/or location, modeled on the STIX 2.1 external_references common property. Cite a CVE, an ATT&CK technique, a STIX bundle/object, a BOM, a runbook, or any vendor artifact. sourceName + externalId is equivalently a URN (urn:<sourceName>:<externalId>), so the by-identity and by-location forms stay interconvertible. A reference is inert context — it overrides nothing and carries only lightweight optional attribution. An embedded document turns a bare reference into a lossless enrichment envelope (the shape hdf enrich writes). Wired onto the results root, Status_Override, and other document types.

FieldTypeRequiredDescription
sourceNamestringyesThe system that names/hosts the artifact (e.g. cve, mitre-attack, stix)
externalIdstringnoIdentifier within sourceName (e.g. CVE-2021-44228)
hrefstringnoURI locating the artifact
descriptionstringnoHuman-readable description
relstringnoOpen relationship token (e.g. investigate, reference)
kindstringnoOpen category token (e.g. threat-intel)
mediaTypestringnoIANA media type of the referenced artifact
checksumChecksumnoIntegrity hash of the referenced artifact
addedByIdentitynoWho attached the reference
addedAtdate-timenoWhen the reference was attached
documentobjectnoLossless embedded copy of the referenced artifact (enrichment envelope)

Integrity Model ​

HDF supports 4 trust levels for tamper detection:

LevelMechanismFields
0None(default)
1Checksumsintegrity on every document root; originalChecksum, resultsChecksum on evaluated baselines
2Amendment chainpreviousChecksum on overrides
3Digital signaturessignature on amendments, evidence packages

Checksum Flow ​

  1. originalChecksum: SHA-256 of the baseline definition file (immutable)
  2. resultsChecksum: SHA-256 of raw results before amendments
  3. previousChecksum: Links each override to the previous override (chain). Each amendment's checksum is recorded by the next one, so an in-place edit is detectable only for an amendment that has a later one chained to it; the final amendment, the document envelope, deleted trailing amendments, and a wholesale recompute of the chain are outside what it can detect. Use signature for non-repudiation. Omitted on the first amendment

Cross-Document References ​

Plan ---- baselineRef ---------> Baseline
Plan ---- systemRef -----------> System
Plan ---- componentRef --------> Component (by UUID)
Results - planRef -------------> Plan
Results - systemRef -----------> System
Results - baselines[] ---------> Baseline (embedded, not referenced)
Results - components[] --------> Component (embedded)
Amendments - systemRef --------> System
Amendments - componentRef -----> Component (by UUID)
Amendments - inheritedFrom ----> Component (by UUID)
Evidence Package - systemRef --> System
Evidence Package - planRef ----> Plan
Evidence Package - componentRef -> Component (by UUID, on Content_Reference)
Results - derivation.seed.uri -> Results (the seed snapshot of a reconciled document)
System - dataFlows[].from/to --> Component (by UUID)
System - controlDesignations --> Component (by UUID, providedBy/inheritedBy)

All *Ref fields are URI-reference strings (relative path, absolute URI, or fragment identifier) unless noted as UUID. Component references use componentId (UUID) for local binding. The baselines[] and components[] relationships in Results are composition — they embed the data rather than referencing external documents.


Schema Compliance ​

All schemas use JSON Schema draft 2020-12 with "unevaluatedProperties": false — documents containing unknown fields are invalid. Tool-specific data should go in the extensions field where available.

Schemas are published at https://mitre.github.io/hdf-libs/schemas/ and distributed as bundled JSON files in the hdf-schema package.

HDF Libraries ​

PackageLanguagePurpose
hdf-schemaTS/GoJSON schemas, generated types, validation helpers
hdf-validatorsTS/GoValidate documents against HDF schemas
hdf-convertersTS/GoConvert 30+ security tool formats to/from HDF
hdf-parsersTS/GoParse and flatten HDF documents
hdf-mappingsTS/GoCCI, CWE, OWASP, NIST framework mappings
hdf-generatorsTS/GoGenerate scanner profile stubs from HDF Baselines
hdf-diffTS/GoCompare HDF documents, produce Comparison documents
hdf-extension-graphTSResolve baseline overlay/extension dependency chains
hdf-utilitiesTSShared utilities: XML parsing, HTML stripping, hashing
hdf-cliGoCLI tool: hdf validate, hdf convert, hdf query, hdf diff

Released under the Apache 2.0 License.