OAG Kenya Findings Corpus · S.I.N.S. Framework™

Kenya Digital
Risk Index

4,675 audit findings extracted from 1,554 public Office of the Auditor-General reports, mapped to the S.I.N.S. Framework™ and aligned to ISO 27005 / NIST CSF. The empirical backbone of RETACH's Cyber Risk Quantification Engine — built from real, cited government audit evidence, not survey data.

4,675Register Rows
562Entities Covered
1,361Reports in Register
9Financial Years
25.6%Findings That Keep Recurring
S.I.N.S. Pillar Distribution

Where the risk
actually sits.

Every finding is tagged multi-label against RETACH's four pillars using curated, word-boundary-matched keyword sets. Most OAG findings are financial/governance in nature — the IT-relevant subset is smaller but consistent, and is exactly the layer RETACH's S.I.N.S. Framework™ is built to fix.

Other (Financial / Governance)95.2% · 4,194 rows
Systems3.0% · 134 rows
Infrastructure1.3% · 59 rows
Security0.8% · 36 rows
Network0.5% · 22 rows
* Network-pillar tags carry a known false-positive risk (keyword echo on non-IT findings) — see Limitations.
S
Systems
ERP hardening, IAM, endpoint control, asset lifecycle.
I
Infrastructure
Server hardening, backup & disaster recovery design.
N
Network
Segmentation, perimeter defense, traffic monitoring.
S
Security
Policy, risk treatment, compliance, awareness culture.
Methodology

Patterns, not predictions.
Proof over guesswork.

We built this the RETACH way: keep it simple, and be honest about what you don't know. The extraction tool looks for clear structural signals in each report — numbered findings, standard section headings — rather than an AI trying to interpret meaning. When a report doesn't give a clean signal, it gets flagged for a person to check by hand, instead of guessing. And before any change went live across all 1,554 reports, it was tested against a set we'd already verified ourselves.

UpdateWhat we fixedResult
v3.6A handful of report titles were unreadable because of a formatting quirk in a few PDFs.6 reports that had failed to process were recovered.
v3.7The tool didn't recognize every way audit reports word their standard legal disclaimers.6 rows that were just boilerplate text, not real findings, were removed.
v3.8Real findings were sometimes getting cut off by the standard disclaimer text sitting in front of them.534 rows fixed (11% of the dataset): 234 real findings were recovered in full, and 300 rows that weren't real findings were correctly removed.

We also map every pillar back to two global standards — ISO 27005 and the NIST Cybersecurity Framework — so the findings aren't just RETACH's opinion. It's how the wider security industry already talks about these risks.

S.I.N.S. PillarISO 27005 DomainNIST CSF Function
SystemsAsset & Vulnerability ManagementProtect
InfrastructurePhysical & Operational SecurityProtect, Recover
NetworkCommunications SecurityProtect, Detect
SecurityRisk Treatment & Organisational ControlsIdentify, Govern
The actuarial way we think about risk: RETACH's Cyber Risk Quantification (CRQ) Engine puts a number on risk the same way insurers do — EAL = Σ(frequency × severity). In plain terms: how often something goes wrong, multiplied by how much it costs when it does, added up across every risk area. This dataset measures the frequency half of that equation — the severity half (what each type of finding actually costs an organisation) is a separate piece we're still building.

The real story behind frequency isn't just "was this flagged before" — it's whether a known problem is left to become normal. We call that the recurrence signal, and we measure it two ways. Auditors themselves flagged 19.8% of findings as an explicit repeat of a prior-year issue. But when we instead traced each entity's own findings across every year we have on file — matching them by what the problem actually is, not just whether the auditor cross-referenced it — 25.6% of all findings turned out to belong to the same unresolved issue recurring across multiple audits. 458 distinct problems, across 243 entities, show up more than once. 341 of those never had a clean year in between — the same issue, audit after audit, unbroken.

Neither number is a finished actuarial coefficient yet. Turning it into one still needs more work: checking it holds up entity by entity, accounting for how recent an issue is, and testing it against reports we haven't looked at yet. We say so here because credibility compounds, and a moat built on inflated claims isn't a moat.

Most recurring threads sit in the general financial/governance category — expected, since OAG audits are financial-statement audits, not IT audits. The slice that matters for RETACH's thesis is the smaller one that ties directly to a S.I.N.S. pillar: 18 of the 458 recurring threads. Here's what that looks like in practice — the same digital-risk issue, named the same way, audit cycle after audit cycle:

EntityS.I.N.S. PillarRecurring issueUnbroken since
Kenya Veterinary BoardInfrastructureLack of Disaster Recovery & Business Continuity Plan2 audits running (2021/22–2022/23)
Tourism Regulatory AuthorityInfrastructureLack of a Disaster Recovery Plan2 audits running (2017/18–2018/19)
Child Welfare Society of KenyaSecurityInadequate IT Governance & Security Policy2 audits running (2018/19–2019/20)
Anti-Counterfeit AuthorityInfrastructureICT Strategy Objectives Not Implemented2 audits running (2018/19–2019/20)
Privatization CommissionSystemsE-Procurement System Not Implemented2 audits running (2021/22–2022/23)
Kenya Universities and Colleges Central Placement ServiceInfrastructureUnsupported ICT Server Procurement2 audits running (2018/19–2019/20)

Smaller n by design — S.I.N.S.-taggable findings are 4.8% of the corpus overall. The pattern (unresolved digital-risk gaps persisting audit after audit) is the point, not the count.

Highest Finding-Count Entities

Top 15 by
finding volume.

Raw finding count, not severity-weighted. High counts partly reflect audit history length and entity complexity, not necessarily worse governance — read alongside the full register.

EntityFindings
National Social Security Fund50
Kenya Pipeline Company Limited44
Western Kenya Rice Mills Limited43
Kenya School of Government42
Tana and Athi Rivers Development Authority41
School Equipment Production Unit39
Water Services Regulatory Board39
Kenya Forest Service38
Kenya Water Institute36
South Nyanza Sugar Company Limited36
Kenya National Shipping Line Limited35
National Museums of Kenya35
Kenyatta International Convention Centre34
Child Welfare Society of Kenya33
Kenyatta National Hospital33
Known Limitations

What this data
doesn't tell you.

RETACH counter-checks its own claims against empirical data before publishing them. These are documented, unresolved gaps — flagged rather than hidden.

Network-pillar false-positive risk

Keyword matches on "network"/"internet" can echo non-IT findings (e.g. a marketing finding mentioning a company's internet presence). Confirmed in manual review at roughly 1-in-3 in a small sample. Network-pillar counts should not be used in public claims without a manual pass.

Scanned-file coverage gap

178 files (11.6% of the corpus) are image-only/scanned PDFs with no extractable text. They are excluded from the register entirely — not represented as empty or "clean" rows — because an unknown is not a zero. OCR recovery is deferred as a phase-2 enhancement.

Year-mismatch drift

17.6% of real findings show a filename-tagged financial year that doesn't match the year stated in the report's own text, versus a 15.0% baseline from earlier sampling — a modest, unexplained +2.6pp drift, not yet root-caused.

Truncation on long findings

Un-numbered findings are stored up to 900 characters; ~14% of these are cut mid-sentence. The underlying extracted text is correct — only the stored excerpt is capped.

The Data Layer Behind The S.I.N.S. Framework™

This is what
execution looks like.

RETACH doesn't sell frameworks — we quantify risk against them, with cited, auditable evidence. This index is the measurement layer behind every Digital Resilience Assessment we deliver.

Talk to RETACH → Visit retach.tech