Predictive Policing Audit Frameworks: What Agencies Must Document to Withstand Oversight
What law enforcement agencies operating predictive policing systems must document, audit, and disclose to satisfy DOJ, NIST, and oversight body scrutiny in 2026.
By IPA-IAC · 7 min · 25 June 2026

The oversight pressure on predictive policing systems has shifted from advocacy to institutional. In 2024, the Department of Justice published a 77-page report on AI in criminal justice that includes explicit documentation and auditing recommendations for agencies operating algorithmic decision-support tools. The Council on Criminal Justice Task Force on AI released a taxonomy in May 2026 that maps governance requirements by risk level. Congress has requested audits of DOJ grant recipients using predictive tools dating back over a decade.
Agencies that deployed these systems between 2015 and 2022 — often with vendor assurances and limited internal oversight infrastructure — are now discovering that the documentation those audits will demand was never created. This article describes what a defensible audit record looks like and what the governance framework must contain to withstand scrutiny from a court, an inspector general, or a legislative oversight body.
What the DOJ December 2024 Report Requires
The DOJ’s December 2024 report on AI and criminal justice is the most direct federal guidance to date on what agencies must maintain. The report’s documentation recommendations are not aspirational — they reflect what DOJ’s own oversight review of agency AI deployments revealed was absent.
The report’s minimum documentation requirements for high-risk AI systems (which includes predictive deployment tools, risk assessment instruments, and facial recognition) are:
System design records: Documentation of what the system does, what data it ingests, what output it produces, and what decisions the output is intended to inform. This includes vendor-provided technical specifications, but not limited to them — agencies should maintain independent descriptions of how the system functions in their specific operational context.
Training data selection and validation records: Documentation of what data was used to train or calibrate the model, who selected it, and what validation was conducted to assess whether the training data reflects enforced patterns that would not be consciously endorsed as policy. This is, per the DOJ report, the most common failure point in retrospective audits.
Testing results: Pre-deployment accuracy and bias testing results, including false positive and false negative rates disaggregated by relevant demographic categories. The DOJ report notes that the appropriate level of testing rigor scales with the consequences of system errors — for predictive patrol tools that affect deployment patterns, demographic disaggregation is not optional.
Ongoing monitoring protocols: Evidence that the agency has a functioning post-deployment monitoring process, not just a documented intention to monitor. This means output logs, defined performance thresholds that trigger review, and a record of what reviews have occurred.
Audit results: Results of any internal or external performance audits, including findings, responses, and system modifications made in response.
The NIST Framework Applied to Predictive Systems
NIST’s AI Risk Management Framework (AI RMF 1.0, released January 2023) provides a governance structure applicable to law enforcement AI that predates most agency deployments and is explicitly referenced in DOJ guidance. The framework’s four core functions — GOVERN, MAP, MEASURE, MANAGE — map directly to predictive policing governance requirements.
GOVERN establishes organizational accountability: who is responsible for AI decisions, what the approval process for deployment is, and how oversight is structured. For predictive policing, this means a named accountability chain from the vendor contract through the patrol commander who acts on system outputs.
MAP identifies risk context: what are the potential harms from this system, who bears those harms, and how do those harms compare to the intended benefits. The CCJ/RAND AI taxonomy recommends that MAP-equivalent analysis for policing tools include explicit equity assessment — who is disproportionately affected when the system errs in a specific direction.
MEASURE creates the evidentiary basis for governance: bias testing, accuracy metrics, performance monitoring, and documentation standards. NIST recommends that audits be conducted by qualified independent bodies — the DOJ report echoes this, noting that vendor self-assessment is insufficient for high-stakes deployments.
MANAGE closes the loop: incident response when system performance degrades, modification protocols, and clear criteria for suspension or discontinuation.
Agencies that have not conducted a NIST AI RMF alignment review of their predictive systems have a governance gap that is visible in any serious audit. The framework is publicly available and widely cited in federal guidance; the absence of documented alignment reads as a policy failure, not a resource constraint.
What Independent Audits Examine
An independent audit of a predictive policing system examines both inputs and outputs. The CCJ/RAND taxonomy recommends this approach explicitly: “Audits should review both the inputs and outputs of systems — the data used to train programs and the decisions they produce.”
Input audit components:
- Data provenance: where did the training or calibration data come from, and what historical enforcement patterns does it encode?
- Proxy variables: does the model use variables that correlate strongly with protected characteristics (race, national origin, religion) without those variables being explicitly included?
- Update cadence: has the model been recalibrated as operational conditions changed, or is it still running on training data from a different enforcement environment?
Output audit components:
- Demographic disparity analysis: do system recommendations result in disproportionate deployment toward specific communities, and is that disparity explained by documented, neutral operational factors or not?
- Override rate analysis: how often do officers act contrary to system recommendations, and what does the pattern of overrides reveal about how the system is actually used?
- Incident correlation: do areas flagged by the system show improved case clearance, reduced crime, or other defined outcome metrics — or is the primary effect increased contact with no outcome improvement?
A 2025 audit of 50 U.S. police departments, cited in multiple policy analyses, found that 68% of predictive policing algorithms in that sample showed racial bias exceeding 22% in patrol deployment recommendations. Only 3% of those departments had formal accountability frameworks to address those disparities. The audit gap is the governance gap.
Congressional Scrutiny and the Grant Accountability Question
Congress formally requested a DOJ audit of all grants issued for predictive policing technology, covering more than a decade of Bureau of Justice Assistance and National Institute of Justice funding. The DOJ’s initial response acknowledged that it “does not have specific records” of how many law enforcement agencies used grants for predictive policing — a documentation failure at the federal level that mirrors what auditors find at the agency level.
The implication for agencies: if federal grant funds were used to purchase, develop, or operate predictive systems, the documentation requirements extend to the grant record. That includes procurement justifications, performance reports submitted to the granting agency, and any compliance certifications about civil rights impact.
Agencies that received BJA or NIJ technology grants between 2014 and 2022 should review their grant documentation for representations about the predictive tools covered, and assess whether their current governance documentation would satisfy a grant compliance review.
Building the Audit-Ready Documentation Set
The minimum documentation set for an agency operating predictive policing tools and seeking to withstand oversight scrutiny has seven components:
System inventory: Every predictive tool in use, including piloted and informally deployed systems, with version, vendor, and function description.
Procurement record: Contract, vendor representations about accuracy and bias performance, and any independent validation documentation provided at procurement.
Deployment authorization: Who approved deployment, on what authority, and what review was conducted prior to deployment.
Training data documentation: Description of training or calibration data, including sources, date range, and any known limitations or biases acknowledged by the vendor.
Pre-deployment testing results: Accuracy and demographic disparity analysis from pre-deployment testing, including the party who conducted the testing.
Ongoing monitoring log: Periodic performance reviews, including metrics, findings, and any system modifications made in response.
Community engagement record: Documentation of any community notification, engagement, or feedback process related to system deployment or operation.
This documentation set does not guarantee favorable audit outcomes. But the absence of it almost always produces adverse findings — not because the system performed poorly, but because the agency cannot demonstrate that it knew how the system performed.
The Oversight Threshold Is Rising
The EPIC comments to DOJ/DHS on law enforcement use of facial recognition, biometric systems, and predictive algorithms in 2024 reflect a policy community that has moved from opposing specific tools to demanding comprehensive governance documentation for an entire category. The standards being applied in oversight reviews in 2026 are substantially higher than those that applied when most predictive policing systems were procured.
Agencies that acquired these systems under lighter governance regimes are not grandfathered against current scrutiny. The documentation and governance infrastructure required today must be built retroactively, against systems that may have been operating for years without it. That work is harder than building governance at the point of procurement — but it is both tractable and necessary before the next audit request arrives.