Borehole · Risk register · free tier
apache/superset
Free reading: not for reliance
Borehole badge for apache/superset
c14ac29 · 5,000 commits · 10,997 files
Issued 2026-09-23

Nothing significant. Questions worth asking about Security & Reliability Posture, Architecture & Codebase — including AI provenance.

35 checks ran across 6 assessable dimensions. None found a significant problem; 3 raised something worth a question. 4 findings, each with evidence. One dimension cannot be assessed without an acquirer.

Degraded run. apache/superset is about 1,139 MB, over the 500 MB free-tier limit, so only the most recent 5,000 commits were read. Signals needing full history are reported as inconclusive rather than clean. Full depth is available on the Report tier.
An earlier reading. A newer reading of this repository exists. This one is kept because its verification code attests to it. Read the newest.
Published under today's rules. The verification code attests to this reading as it was stored. This page differs from it as follows.
  • Words about people are today's: bands rather than exact shares, dates or day counts, and what the history measured rather than why.
  • 1 finding that needed more history than this survey read is no longer published.
  • 1 finding from checks Borehole's calibration found wrong more often than right is no longer published.
strong adequate significant findings not assessable from a repository
Provenance disclosure does not affect any rating
78% of added lines arrived in commits declaring an AI agent as co-author
The surviving-lines measurement was not run on this repository.
1,707/5,000 commits declaring an agent
confidence: moderate
  • 75% of non-merge commits carry any trailer at all
  • no widespread author/committer date divergence
  • the declared share is neither decisive nor negligible, so the true figure is more sensitive to undeclared work
324,888 lines excluded as lockfiles, vendored or generated.
This is a lower bound. It counts what the repository declares. Undeclared AI authorship is invisible here, and a single rebase removes the trailers this rests on. Absence of a declaration is absence of disclosure, not absence of AI authorship.

Architecture & Codebase — including AI provenance

01 · adequate · ratings ↑
01.6

tests/unit_tests/sql/parse_tests.py is 6,683 lines, 63× this codebase's median file

adequate

The median source file here is 105 lines. 201 files exceed 1,050 lines, the largest being 6,683. Size is not a defect on its own — some problems genuinely live in one place — but these are where merge conflicts, review fatigue and single-owner knowledge concentrate, and they are the first thing that slows an inheriting engineer down.

Evidence
line-count tests/unit_tests/sql/parse_tests.py:1 — 6,683 lines (63× median)
line-count superset/security/manager.py:1 — 6,362 lines (60× median)
line-count tests/integration_tests/dashboards/api_tests.py:1 — 6,020 lines (57× median)
line-count tests/unit_tests/commands/report/execute_test.py:1 — 5,902 lines (56× median)
line-count superset/models/helpers.py:1 — 5,802 lines (55× median)
Respond to this finding

Your response is published beside the finding, not instead of it. A finding you can argue with is the point — and a wrong one is calibration data we want.

An owner's response is removed only when it is unlawful, impersonates someone, publishes personal data or a credential, or is not about the finding. Disagreeing with the report is never a reason. To report a response, or to reach us about one, write to takedown@borehole.dev.

Engineering Process & SDLC Maturity

02 · strong · ratings ↑

No finding. 7 checks ran and none fired; the rejection log below says why each did not.

Organization & Key-Person Risk

03 · strong · ratings ↑

No finding. 5 checks ran and none fired; the rejection log below says why each did not.

Product & Engineering Maturity

04 · adequate · ratings ↑
04.3

142 direct dependencies in a single manifest

adequate

superset-frontend/package.json declares 142 direct dependencies. Each one is a maintenance obligation, a supply-chain surface and a potential licence question that transfers with the asset. The count says nothing about whether any given dependency was a good choice; it says how many such choices a buyer inherits.

Evidence
direct-dependencies superset-frontend/package.json — 142 declared
direct-dependencies superset-frontend/packages/superset-ui-core/package.json — 45 declared
direct-dependencies docs/package.json — 39 declared
direct-dependencies superset-frontend/plugins/preset-chart-deckgl/package.json — 30 declared
Respond to this finding

Your response is published beside the finding, not instead of it. A finding you can argue with is the point — and a wrong one is calibration data we want.

An owner's response is removed only when it is unlawful, impersonates someone, publishes personal data or a credential, or is not about the finding. Disagreeing with the report is never a reason. To report a response, or to reach us about one, write to takedown@borehole.dev.

Security & Reliability Posture

05 · adequate · ratings ↑
05.6

1 Dockerfile(s) run the application as root

adequate

No USER instruction switches away from root, so the process runs with full privileges inside the container. It is a one-line fix and an ordinary expectation, which is why its absence is worth asking about: it usually says more about whether anyone has reviewed the deployment than about this container specifically.

Evidence
no-USER-instruction .devcontainer/Dockerfile — runs as root
Respond to this finding

Your response is published beside the finding, not instead of it. A finding you can argue with is the point — and a wrong one is calibration data we want.

An owner's response is removed only when it is unlawful, impersonates someone, publishes personal data or a credential, or is not about the finding. Disagreeing with the report is never a reason. To report a response, or to reach us about one, write to takedown@borehole.dev.

05.9

2 vendored components with no licence

adequate

Third-party code is committed into this repository without any statement of the terms it arrived under. Unknown terms are harder to deal with than inconvenient ones: an awkward licence can be complied with, an absent one cannot. Each of these needs its origin established before the code changes hands.

Evidence
superset-frontend/plugins/plugin-chart-calendar/src/vendor/cal-heatmap.ts superset-frontend/plugins/plugin-chart-calendar/src/vendor/cal-heatmap.ts — 1 tracked files, no licence file
superset-frontend/plugins/plugin-chart-parallel-coordinates/src/vendor/parcoords superset-frontend/plugins/plugin-chart-parallel-coordinates/src/vendor/parcoords — 2 tracked files, no licence file
Respond to this finding

Your response is published beside the finding, not instead of it. A finding you can argue with is the point — and a wrong one is calibration data we want.

An owner's response is removed only when it is unlawful, impersonates someone, publishes personal data or a credential, or is not about the finding. Disagreeing with the report is never a reason. To report a response, or to reach us about one, write to takedown@borehole.dev.

Observability & Operations

06 · strong · ratings ↑

No finding. 3 checks ran and none fired; the rejection log below says why each did not.

Fit with the Acquirer

07 · not assessable · ratings ↑

Not rated: structurally not assessable from a repository. What a person adds here.

Check it on the running system

2 checks against the running system follow from the findings above: what a buyer's engineer will try, how to run it first, and what to have ready. A paid survey of this repository includes them. See the tiers.

A paid survey of a public repository is also reviewed: a language model reads each finding against the code and withdraws only findings it judges the code contradicts, each listed with its reason. The review is a language model's reading of the findings against the code, and it can be wrong. It never adds a finding, never raises a band, and never removes a finding from a proven check or an advisory that affects the surveyed version.

What a person adds

The survey reads what the repository can show. The code cannot settle these questions, and on them this report states nothing. Each says where its answer lives: your own engineer, the company, or an assessor you choose can settle it. Questions marked measured open with something this repository shows.

Engineering Process & SDLC Maturity

02 · high coverage
  1. Incident response practice is not visible in a repository.

Organization & Key-Person Risk

03 · partial coverage
  1. Hiring pipeline, bench depth and retention are not visible in a repository.
  2. Whether the dominant contributor is part of the transaction.
  3. Seniority, levelling and who actually mentors whom.
  4. Whether the team wants to stay after the deal closes.

Product & Engineering Maturity

04 · minimal coverage
  1. Roadmap credibility requires the data room.
  2. Whether technical debt is acknowledged and funded.

Security & Reliability Posture

05 · partial coverage
  1. Incident history and failure-mode analysis are not in the repository.

Observability & Operations

06 · partial coverage
  1. Whether anyone watches the telemetry that exists.

Fit with the Acquirer

07 · none coverage

Not rated: structurally not assessable from a repository.

  1. This is a Python and TypeScript codebase. Does your engineering organisation already run Python and TypeScript in production, or would this be the first?
  2. 423 people have committed here. Which of them are you acquiring, and which are you replacing?
  3. The bulk of this system sits in superset/, superset-frontend/, tests/. Which of your existing systems would each of those have to absorb or replace?
  4. 78% of added lines came through AI-declared commits. Does your organisation have a policy on inheriting and maintaining that code, and does your engineering team's tooling match how it was produced?
  5. This deploys as a container or managed service. Does that match your deployment and on-call model, or does it need rehosting first?
  6. What is your tolerance for the findings above remaining unaddressed for the first two quarters after close, given everything else that will compete for this team's time?

A human assessment is Borehole's own paid service, €12,000+ (total price is determined per engagement; travel billed at cost). It is not independent of this report: the same business produced both. It is a separate contract, with its own scope, terms and reliance.

Embed the badge

Tied to this commit and this date, and it fades after 90 days — an honest badge is one that expires. A repository with significant findings gets a neutral "read" badge rather than a red one: the library is where the range is visible, not someone else's README.

Borehole badge for apache/superset
[![Borehole](/badge/apache/superset.svg "Borehole free reading: not for reliance. See the full report.")](/r/apache/superset/c14ac29)

If you embed it: link it to the report, remove it if the report is delisted, and do not present it as a certification or an endorsement.

Compare with an earlier read
Was this reading useful?
Rejection log — 31 signals that ran and did not fire

Appendix: rejection log — 31 signals that ran and did not fire

Published on purpose. A signal that looked and decided not to fire has told you something, and this is the calibration data behind every threshold above.

Looked, and found nothing — 28

SignalWhat it found
04.hotspot_concentration change is spread across the codebase rather than pooling in a few filesfiles 13,167 · top5 share 0.068
04.debt_markers debt markers are sparsemarkers 730 · per kloc 0.42 · lines read 1,754,726
04.test_growth test effort kept pace with code writtenlate test per code 0.402 · early test per code 0.364
04.churn_trend rework is not rising against new worklate delete per insert 0.339 · early delete per insert 0.118
06.telemetry_presence a telemetry or error-reporting library is referencednote presence only; whether anyone watches it is not visible here · matches 40
06.structured_logging a structured logging library is referencedmatches 9
06.healthcheck_presence a health or readiness check is defined
03.contributor_concentration no single contributor dominates the historyauthors 423 · top share a quarter or more · bus factor 2
03.knowledge_silos every substantial directory has been touched by more than one personauthors 423 · directories 15
03.drive_by_contributors a healthy share of contributors committed more than onceshare 0.686 · authors 423 · one commit authors 290
03.contributor_attrition the dominant contributor is still activeauthors 423 · top share a quarter or more · top contributor last committed within the last three months
02.merge_discipline a meaningful share of change arrives through a reviewable pathmerges 0 · commits 5,000 · reviewed share 0.999
02.ci_presence CI configuration is present in the working treecount 55
02.test_presence tests are present in a normal proportion to sourcetests 2,629 · source 6,905 · test ratio 0.381
02.release_cadence tags are presenttags 25
02.conventional_commits commit subjects generally carry contentsampled 5,000 · low content share 0.0
02.branch_protection_hints process scaffolding is present in the repository
02.ci_runs_tests CI invokes a test runnerconfigs 55 · with tests 7
01.message_uniformity commit subjects vary in wording and lengthlength stdev 16.06 · repeat share 0.0
01.revision_depth most files were revisited after they were first writtenfiles seen 13,167 · write once share 0.464
01.duplicated_blocks little shipping code is repeated verbatim across filesshare 0.0128 · product files 5,012 · distinct blocks 590,541
05.committed_secrets no credential-shaped strings outside test and fixture pathsnote the working tree only; history is scanned at --deep · patterns 7 · files read 8,731
05.security_guardrails at least one standard guard rail is configured
05.dependency_pinning a lockfile is committed, so builds are reproduciblelockfile True
05.license_present a licence file is presentfound True
05.secrets_in_history no credential-shaped string was ever added outside test and fixture pathspatterns 7 · truncated False · added lines 1,403,210
05.license_copyleft_inherited every licence found in vendored code is permissive or weak-copyleftvendored paths 3
05.license_grant_unclear the licence is recognisable as Apache-2.0class permissive · license Apache-2.0

Could not conclude, or did not apply — 3

Not a clean result. Each says why it could not answer here.

SignalWhy
03.active_contributorshistory too short to distinguish active from historical contributors
05.license_manifest_conflictno packaging manifest states a licence
01.initial_dumpnot assessed: this check needs the whole history, and only the most recent 5,000 commits were read
BOREHOLE READING NOT A CERTIFICATION 69DY-71QT-Q1JM 2026-09-23 · C14AC29 BOREHOLE.DEV/VERIFY

Anyone can check this report against the authoritative record — click the stamp, or enter 69DY-71QT-Q1JM at borehole.dev/verify.

The code commits to the 36 checks this run performed, the bands they produced, and the commit they were read at. It does not cover checks written since.

Put the mark in your README

The image is wrapped in a link to the verification page, so a reader can check it rather than take it on trust. If you embed it: link it to the report, remove it if the report is delisted, and do not present it as a certification or an endorsement.

[![Borehole reading, not a certification](https://borehole.dev/badge/apache/superset/c14ac29.svg "Borehole free reading: not for reliance. See the full report.")](https://borehole.dev/verify/69DY71QTQ1JM?via=badge)

Borehole reading 69DY-71QT-Q1JM, not a certification

What this scan did with the repository — 5 steps

Written by the steps themselves as they ran, not described afterwards. A hand-written account of a data flow drifts the first time somebody changes the flow and not the account.

AtStepDetail
0.4s read repository metadata from the GitHub API size mb: 1138.5authenticated: False
38.1s cloned the repository to a temporary directory depth: 5,000 commits
60.7s read the git history commits: 5000tracked files: 10997
122.0s ran the deterministic checks blame pass: Falseread working tree: True
122.1s deleting the working copy

Total 122.1s ·source content sent off this machine: none · calls to a language model: 0 · the clone was deleted when the survey ended

An automated reading of a repository's code and its history up to one commit, with public advisory records. It may be wrong, it is not legal, financial, investment or security advice, and it does not replace diligence by a qualified person. It is provided for information: nobody may rely on it (Borehole's terms, section 3).

BOREHOLE 0.1.0.DEV0 DETERMINISTIC LAYER · NO MODEL CALLSREAD ON BOREHOLE'S SERVER · CLONE DELETED TAKEDOWNCONTRIBUTED TO THIS REPOSITORY? YOUR RIGHTS