Repository navigation
feat(metrics): add COCO keypoint mAP (OKS) for sv.KeyPoints - #2687
Merged
Merged
Conversation
8rulerstar
added a commit
to 8rulerstar/supervision
that referenced
this pull request
Oct 7, 2026
2 tasks done
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## develop #2687 +/- ##
========================================
Coverage 93% 93%
========================================
Files 81 82 +1
Lines 12238 12526 +288
========================================
+ Hits 11360 11674 +314
+ Misses 878 852 -26 🚀 New features to boost your workflow:
|
Contributor
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Valid three-coordinate KeyPoints can crash or miscompute OKS, while infinite sigmas can make unrelated poses perfect matches.
Review effort: Balanced
Findings: 1
Open (7)
Handle 3-column keypoints as planar coordinates with visibility · New Reject non-finite sigma values · New Document zero-area OKS behavior accurately · New Correct zero-area OKS behavior description · New Clarify float32 bias removal in the spacing comment · New Add semantic IDs to compound parameter sets · New Move test into the MeanAveragePrecision test class · New
What changed in this PR
Adds COCO-compatible keypoint mAP evaluation for sv.KeyPoints using Object Keypoint Similarity.
Changes:
- Introduces keypoint mAP metrics, OKS matching, and public exports.
- Adds extensive parity and regression tests.
- Adds API documentation, navigation, and changelog entries.
| File | Description |
|---|---|
src/supervision/detection/utils/iou_and_nms.py |
Implements OKS calculation and COCO sigmas. |
src/supervision/metrics/keypoint_mean_average_precision.py |
Implements keypoint mAP evaluation and result reporting. |
src/supervision/metrics/mean_average_precision.py |
Shares reporting helpers and updates precision/max-detection handling. |
src/supervision/metrics/__init__.py |
Exports the new metric API. |
tests/metrics/test_keypoint_mean_average_precision.py |
Tests OKS, COCO parity, categories, and reporting. |
tests/metrics/test_mean_average_precision.py |
Adds epsilon and max-detection regressions. |
docs/metrics/keypoint_mean_average_precision.md |
Documents the public API. |
mkdocs.yml |
Adds the documentation page to navigation. |
docs/changelog.md |
Records the feature and precision fix. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
8rulerstar
added a commit
to 8rulerstar/supervision
that referenced
this pull request
Oct 7, 2026
8rulerstar
force-pushed
the
feat/keypoint-map
branch
from
October 7, 2026 07:13
572f275 to
800d71f
Compare
Add sv.metrics.KeypointMeanAveragePrecision, the COCO keypoint mAP over sv.KeyPoints predictions and targets, plus sv.keypoint_oks_batch and sv.COCO_KEYPOINT_SIGMAS. The metric subclasses the existing COCOEvaluator and only replaces the IoU computation with OKS, so matching, accumulation and 101-point interpolation are shared with MeanAveragePrecision. The accumulator now reads the largest max_dets instead of the literal 100, which keeps box, mask and OBB results unchanged and allows the COCO keypoint maxDets of 20. Scores match pycocotools 2.0.11 COCOeval(..., "keypoints") on synthetic data within 1e-7.
…tools The precision `tp / (fp + tp + EPS)` used the float32 epsilon (~1.2e-7) where pycocotools uses `np.spacing(1)` (~2.2e-16), so a perfect precision became 0.99999988 and one perfectly predicted box scored mAP@50:95 0.9999998807907104 instead of 1.0. Use `np.spacing(1)`.
8rulerstar
force-pushed
the
feat/keypoint-map
branch
from
October 7, 2026 07:47
800d71f to
ef33c8c
Compare
Changes: - Mask non-finite keypoint coordinate deltas before adding similarity, so each contributes zero without contaminating finite points. - Add a mixed-NaN regression proving visible finite points retain expected contribution. Impact: - Prevents NaN coordinates in predictions from turning a valid OKS result into NaN. - Keeps visible-keypoint denominator behavior unchanged while preserving finite keypoint scores. Verification: - Full suite via .venv/bin/pytest --cov=supervision: 4,811 passed, 61 skipped, 39 warnings. - pre-commit run --all-files: passed. Residual limits: - Codex Rig workflow remediation remains outside this repository commit because its worktree mixes those changes with unrelated uncommitted work. - OpenCV is unavailable; tests used the package’s NumPy fallback. --- Co-authored-by: Codex <codex@openai.com>
Contributor
There was a problem hiding this comment.
🟡 Changes recommended
NaN handling and its documentation do not match the claimed pycocotools behavior.
3 open findings
7 resolved since last review
Handle 3-column keypoints as planar coordinates with visibility Reject non-finite sigma values Move test into the MeanAveragePrecision test class Add semantic IDs to compound parameter sets Clarify float32 bias removal in the spacing comment Correct zero-area OKS behavior description Document zero-area OKS behavior accurately
🧠 Review effort: Balanced
- fix(metrics): treat non-finite target keypoints as unlabelled - fix(metrics): validate keypoint mAP area and sigmas upfront - docs(metrics): document _compute_iou as the COCO evaluator hook - refactor(metrics): move mAP result helpers to metrics core - fix(metrics): keep class 0 for unlabelled agnostic keypoint mAP - fix(metrics): reject keypoint skeletons with zero keypoints [resolve group] PR roboflow#2687 — items 12 13 14 15 16 32 --- Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com> Co-authored-by: Codex <codex@openai.com>
- test(metrics): add pycocotools parity generator and state the ~1e-7 tolerance - test(metrics): pin keypoint mAP fallback target area semantics - test(metrics): cover ignore-region and crowd semantics in keypoint mAP - test(metrics): pin keypoint mAP maxDets cap and missing-confidence scoring [resolve group] PR roboflow#2687 — items 17 18 19 20 --- Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com> Co-authored-by: Codex <codex@openai.com>
- perf(metrics): skip small area range in keypoint mAP evaluation - perf(metrics): vectorize per-image keypoint mAP prep - refactor(metrics): add public xyxy and iscrowd data keys to config [resolve group] PR roboflow#2687 — items 22 23 27 --- Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com> Co-authored-by: Codex <codex@openai.com>
- docs(metrics): state COCO v=1 and v=2 map to visible keypoints - docs(metrics): clarify per-class prediction cap and bbox wording in keypoint mAP [resolve group] PR roboflow#2687 — items 24 25 --- Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com> Co-authored-by: Codex <codex@openai.com>
[resolve group] PR roboflow#2687 — items 21 --- Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com> Co-authored-by: Codex <codex@openai.com>
[resolve group] PR roboflow#2687 — items 28 --- Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com> Co-authored-by: Codex <codex@openai.com>
- `_pycocotools_summarize` looks up the size index by area range via `_OBJECT_SIZE_AREA_RANGES`, not `ObjectSize` enum position - drop private `_XYXY_DATA_FIELD`/`_ISCROWD_DATA_FIELD` aliases; use `XYXY_DATA_FIELD`/`ISCROWD_DATA_FIELD` from `config.py` directly - docs: a class whose only targets are skipped (no keypoints, no `xyxy`) scores `-1` --- Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com>
- `_KEYPOINT_MAX_DETECTIONS` and `_COCO_KEYPOINT_SIGMAS`: replace the floating string literal after the assignment with a `#:` comment above it --- Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com>
- rename `tests/metrics/generate_keypoint_map_parity.py` to `_generate_keypoint_map_parity.py` - update the `python -m tests.metrics._generate_keypoint_map_parity` command in its docstring and in the parity test comment --- Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com>
- `mkdocs.yml`: `Keypoint mAP` nav label becomes `mAP (KeyPoints)` --- Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com>
- `mkdocs.yml`: bare `mAP` nav label becomes `mAP (Detections)`, pairing with `mAP (KeyPoints)` --- Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com>
- `docs/metrics/mean_average_precision.md`: intro covering both `MeanAveragePrecision` (Detections, IoU) and `KeyPointMeanAveragePrecision` (KeyPoints, OKS); new `## Pose estimation` section with the keypoint usage notes and example; `KeyPointMeanAveragePrecision` and `KeyPointMeanAveragePrecisionResult` API blocks appended - remove `docs/metrics/keypoint_mean_average_precision.md` - `mkdocs.yml`: single `mAP` nav entry replaces `mAP (Detections)` and `mAP (KeyPoints)` --- Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com>
Borda
approved these changes
Oct 9, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



Description
Closes #2686.
Thank you for maintaining supervision! This follows up on the feature proposal in #2686.
Adds
sv.metrics.KeypointMeanAveragePrecision, the COCO keypoint mAP. It reusesCOCOEvaluatorand replaces only the IoU step with Object Keypoint Similarity, ported frompycocotools'computeOkswith the keypoint parameters (maxDets=20, medium/large areas, COCO sigmas). Ignore regions for targets without visible keypoints and crowd targets follow COCO.The
fix(metrics)commit is a small related fix: the shared precision guard usednp.finfo(np.float32).eps, whilepycocotoolsusesnp.spacing(1), so a perfect box prediction scored 0.99999988 instead of 1.0. It is its own commit so it can be split into a separate PR if you prefer.Motivation and Context
There is no pose metric in
sv.metricsyet; evaluatingsv.KeyPointsmeans converting to COCO JSON and runningpycocotools. Details in #2686. I had no answers to the open questions there yet, so I went with the defaults proposed in the issue and am happy to change any of them:_keypoint_oks_batch,_COCO_KEYPOINT_SIGMAS,_KEYPOINT_MAX_DETECTIONS).data["area"],data["xyxy"]anddata["iscrowd"]; predictions readdata["xyxy"]for the size bucket.pycocotoolscomparison is a script in this description, not a test dependency.Changes Made
metrics/keypoint_mean_average_precision.py:KeypointMeanAveragePrecisionandKeypointMeanAveragePrecisionResult(plot(),to_pandas(), printable COCO-style summary).detection/utils/iou_and_nms.py: private_keypoint_oks_batch, including the COCO expanded-box branch for targets without visible keypoints.metrics/mean_average_precision.py:max_detsinstead of the literal100. Box, mask and OBB results are unchanged with the default settings; a regression test covers it.EPS = np.spacing(1).Docs page, mkdocs entry, changelog entries.
Review follow-up: keypoints given as
(N, K, 3)use only the planarxy[..., :2];sigmasmust be positive and finite; docs describe NaN keypoints and zero target areas precisely; more tests for result helpers and edge cases.Testing
Compared with
pycocotools2.0.11COCOeval(..., "keypoints"), all five summary values (AP, AP50, AP75, APm, APl):yolo11n-posepredictions, 853 targets incl. 19 crowd and 267 without visible keypointsWithout target boxes, targets with no visible keypoints are skipped and the COCO subset differs by up to 8.6e-3; this is documented.
uv run pytest tests/metrics tests/detection: 2673 passed, 1 skipped. Doctests of the changed modules pass.pre-commiton changed files passes (incl. mypy).The epsilon test fails on
develop(0.9999998807907104 == 1.0) and passes with the fix. On random data, box, mask and OBB scores move up by at most 5e-8.Added/updated tests, and the full suite passes locally
Updated docs (docstrings / mkdocs entry) for new or changed public API
Added a changelog entry in
docs/changelog.mdunderUnreleased(skip for lint/type/format-only or pure doc changes)Additional Notes
keypoint visibleis boolean insv.KeyPoints, so COCOv=1andv=2are both treated as labelled, as inpycocotools.Thank you for taking the time to review this. Any feedback is very welcome, and I'm glad to adjust the API, naming or structure to whatever fits supervision best.