feat(nativemem): categorized native-memory accounting — first cut#669
feat(nativemem): categorized native-memory accounting — first cut#669rkennke wants to merge 8 commits into
Conversation
CI Test ResultsRun: #29833525967 | Commit:
Status Overview
Legend: ✅ passed | ❌ failed | ⚪ skipped | 🚫 cancelled Summary: Total: 32 | Passed: 32 | Failed: 0 Updated: 2026-07-21 13:31:57 UTC |
Benchmark Results (commit e095203)Pipeline: https://gitlab.ddbuild.io/DataDog/apm-reliability/benchmarking-platform/-/pipelines/125289326 Commit:
|
| Benchmark | JDK | Latest | Dev | Δ (dev vs latest) | Issues L/D |
|---|---|---|---|---|---|
| akka-uct | 21 | ✅ 10287 ms (21 iters) | ✅ 10260 ms (21 iters) | ≈ -0.3% (±10.9%) | — / — |
| akka-uct | 25 | ✅ 8867 ms (24 iters) | ✅ 8805 ms (24 iters) | ≈ -0.7% (±9.8%) | — / — |
| finagle-chirper | 21 | ✅ 5933 ms (33 iters) | ✅ 5935 ms (33 iters) | ≈ +0% (±24.9%) | |
| finagle-chirper | 25 | ✅ 5475 ms (36 iters) | ✅ 5453 ms (36 iters) | ≈ -0.4% (±24.5%) | |
| fj-kmeans | 21 | ✅ 2648 ms (71 iters) | ✅ 2700 ms (69 iters) | ≈ +2% (±2.7%) | — / — |
| fj-kmeans | 25 | ✅ 2825 ms (66 iters) | ✅ 2836 ms (66 iters) | ≈ +0.4% (±2.6%) | — / — |
| future-genetic | 21 | ✅ 2121 ms (88 iters) | ✅ 2045 ms (90 iters) | 🟢 -3.6% | — / — |
| future-genetic | 25 | ✅ 1976 ms (94 iters) | ✅ 2086 ms (89 iters) | 🔴 +5.6% | — / — |
| naive-bayes | 21 | ✅ 1230 ms (139 iters) | ✅ 1259 ms (137 iters) | ≈ +2.4% (±32.9%) | — / — |
| naive-bayes | 25 | ✅ 1015 ms (168 iters) | ✅ 1011 ms (169 iters) | ≈ -0.4% (±31.6%) | — / — |
| reactors | 21 | ✅ 15811 ms (16 iters) | ✅ 16445 ms (15 iters) | ≈ +4% (±7.6%) | — / — |
| reactors | 25 | ✅ 18243 ms (15 iters) | ✅ 18764 ms (15 iters) | ≈ +2.9% (±4.1%) | — / — |
Internal counter details (ddprof)
ddprof internal counters, latest / dev (✅ = 0, · = unavailable):
| Benchmark | JDK | Dropped rec | Dropped jvmti | Dropped trace | Skipped WC | AGCT fail | Unwind fail |
|---|---|---|---|---|---|---|---|
| akka-uct | 21 | ✅ / ✅ | ✅ / ✅ | 3 / 1 | 2094 / 1933 | ✅ / ✅ | ✅ / ✅ |
| akka-uct | 25 | ✅ / ✅ | ✅ / ✅ | 1 / 3 | 2216 / 2195 | ✅ / ✅ | ✅ / ✅ |
| finagle-chirper | 21 | ✅ / ✅ | ✅ / ✅ | 4 / 2 | 8790 / 8777 | ✅ / ✅ | ✅ / ✅ |
| finagle-chirper | 25 | ✅ / ✅ | ✅ / ✅ | 1 / 3 | 8165 / 8254 | ✅ / ✅ | ✅ / ✅ |
| fj-kmeans | 21 | ✅ / ✅ | ✅ / ✅ | 2 / 2 | 1283 / 1240 | ✅ / ✅ | ✅ / ✅ |
| fj-kmeans | 25 | ✅ / ✅ | ✅ / ✅ | 1 / 2 | 1289 / 1281 | ✅ / ✅ | ✅ / ✅ |
| future-genetic | 21 | ✅ / ✅ | ✅ / ✅ | ✅ / 2 | 2935 / 2926 | ✅ / ✅ | ✅ / ✅ |
| future-genetic | 25 | ✅ / ✅ | ✅ / ✅ | 1 / 2 | 2884 / 2927 | ✅ / ✅ | ✅ / ✅ |
| naive-bayes | 21 | ✅ / ✅ | ✅ / ✅ | 4 / 2 | 3486 / 3580 | ✅ / ✅ | ✅ / ✅ |
| naive-bayes | 25 | ✅ / ✅ | ✅ / ✅ | 5 / ✅ | 3448 / 3533 | ✅ / ✅ | ✅ / ✅ |
| reactors | 21 | ✅ / ✅ | ✅ / ✅ | 1 / ✅ | 1512 / 1844 | ✅ / ✅ | ✅ / ✅ |
| reactors | 25 | ✅ / ✅ | ✅ / ✅ | 1 / 1 | 1866 / 1889 | ✅ / ✅ | ✅ / ✅ |
This comment has been minimized.
This comment has been minimized.
Reliability & Chaos Results✅ All reliability & chaos checks passed Pipeline: https://gitlab.ddbuild.io/DataDog/java-profiler/-/pipelines/125658657 |
Benchmark Results (commit 62b6f2c)Pipeline: https://gitlab.ddbuild.io/DataDog/apm-reliability/benchmarking-platform/-/pipelines/125313150 Commit:
|
| Benchmark | JDK | Latest | Dev | Δ (dev vs latest) | Issues L/D |
|---|---|---|---|---|---|
| akka-uct | 21 | ✅ 10303 ms (21 iters) | ✅ 10356 ms (21 iters) | ≈ +0.5% (±10.7%) | — / — |
| akka-uct | 25 | ✅ 9028 ms (24 iters) | ✅ 8895 ms (24 iters) | ≈ -1.5% (±9.8%) | — / — |
| finagle-chirper | 21 | ✅ 5921 ms (33 iters) | ✅ 5936 ms (33 iters) | ≈ +0.3% (±25.2%) | |
| finagle-chirper | 25 | ✅ 5537 ms (36 iters) | ✅ 5535 ms (36 iters) | ≈ -0% (±24.4%) | |
| fj-kmeans | 21 | ✅ 2705 ms (70 iters) | ✅ 2667 ms (69 iters) | ≈ -1.4% (±2.6%) | — / — |
| fj-kmeans | 25 | ✅ 2829 ms (66 iters) | ✅ 2824 ms (66 iters) | ≈ -0.2% (±2.6%) | — / — |
| future-genetic | 21 | ✅ 2122 ms (88 iters) | ✅ 2066 ms (90 iters) | 🟢 -2.6% | — / — |
| future-genetic | 25 | ✅ 2071 ms (90 iters) | ✅ 1960 ms (94 iters) | 🟢 -5.4% | — / — |
| naive-bayes | 25 | ✅ 1022 ms (167 iters) | ✅ 1012 ms (169 iters) | ≈ -1% (±31.6%) | — / — |
| reactors | 21 | ✅ 16191 ms (15 iters) | ✅ 15952 ms (15 iters) | ≈ -1.5% (±6.7%) | — / — |
| reactors | 25 | ✅ 18581 ms (15 iters) | ✅ 18426 ms (15 iters) | ≈ -0.8% (±3.8%) | — / — |
Internal counter details (ddprof)
ddprof internal counters, latest / dev (✅ = 0, · = unavailable):
| Benchmark | JDK | Dropped rec | Dropped jvmti | Dropped trace | Skipped WC | AGCT fail | Unwind fail |
|---|---|---|---|---|---|---|---|
| akka-uct | 21 | ✅ / ✅ | ✅ / ✅ | 2 / ✅ | 1975 / 2029 | ✅ / ✅ | ✅ / ✅ |
| akka-uct | 25 | ✅ / ✅ | ✅ / ✅ | ✅ / ✅ | 2406 / 2283 | ✅ / ✅ | ✅ / ✅ |
| finagle-chirper | 21 | ✅ / ✅ | ✅ / ✅ | 7 / 7 | 8749 / 8443 | ✅ / ✅ | ✅ / ✅ |
| finagle-chirper | 25 | ✅ / ✅ | ✅ / ✅ | 1 / 2 | 8539 / 8755 | ✅ / ✅ | ✅ / ✅ |
| fj-kmeans | 21 | ✅ / ✅ | ✅ / ✅ | 5 / 3 | 1305 / 1244 | ✅ / ✅ | ✅ / ✅ |
| fj-kmeans | 25 | ✅ / ✅ | ✅ / ✅ | 4 / 1 | 1284 / 1292 | ✅ / ✅ | ✅ / ✅ |
| future-genetic | 21 | ✅ / ✅ | ✅ / ✅ | ✅ / 1 | 3065 / 2961 | ✅ / ✅ | ✅ / ✅ |
| future-genetic | 25 | ✅ / ✅ | ✅ / ✅ | 2 / 4 | 2938 / 2818 | ✅ / ✅ | ✅ / ✅ |
| naive-bayes | 25 | ✅ / ✅ | ✅ / ✅ | 5 / 4 | 3483 / 3485 | ✅ / ✅ | ✅ / ✅ |
| reactors | 21 | ✅ / ✅ | ✅ / ✅ | 1 / ✅ | 1495 / 1554 | ✅ / ✅ | ✅ / ✅ |
| reactors | 25 | ✅ / ✅ | ✅ / ✅ | 1 / ✅ | 1874 / 1984 | ✅ / ✅ | ✅ / ✅ |
Benchmark Results (commit bdf9f55)Pipeline: https://gitlab.ddbuild.io/DataDog/apm-reliability/benchmarking-platform/-/pipelines/125648678 Commit: ✅ Within expected boundariesNo significant runtime deltas (all within run-to-run noise) and no internal-counter outliers. Runtime details (per benchmark × JDK)
Internal counter details (ddprof)ddprof internal counters, latest / dev (✅ = 0, · = unavailable):
|
Benchmark Results (commit f1a84c6)Pipeline: https://gitlab.ddbuild.io/DataDog/apm-reliability/benchmarking-platform/-/pipelines/125658666 Commit:
|
| Benchmark | JDK | Latest | Dev | Δ (dev vs latest) | Issues L/D |
|---|---|---|---|---|---|
| akka-uct | 21 | ✅ 10309 ms (21 iters) | ✅ 10323 ms (21 iters) | ≈ +0.1% (±10.8%) | — / — |
| akka-uct | 25 | ✅ 8961 ms (24 iters) | ✅ 8884 ms (24 iters) | ≈ -0.9% (±9.8%) | — / — |
| finagle-chirper | 21 | ✅ 6042 ms (33 iters) | ✅ 5955 ms (33 iters) | ≈ -1.4% (±25.3%) | |
| finagle-chirper | 25 | ✅ 5468 ms (36 iters) | ✅ 5531 ms (36 iters) | ≈ +1.2% (±25%) | |
| fj-kmeans | 21 | ✅ 2649 ms (70 iters) | ✅ 2695 ms (70 iters) | ≈ +1.7% (±2.7%) | — / — |
| fj-kmeans | 25 | ✅ 2800 ms (67 iters) | ✅ 2807 ms (66 iters) | ≈ +0.3% (±2.6%) | — / — |
| future-genetic | 21 | ✅ 2141 ms (87 iters) | ✅ 2039 ms (90 iters) | 🟢 -4.8% | — / — |
| future-genetic | 25 | ✅ 2088 ms (89 iters) | ✅ 2127 ms (87 iters) | ≈ +1.9% (±2.5%) | — / — |
| naive-bayes | 21 | ✅ 1227 ms (139 iters) | ✅ 1280 ms (134 iters) | ≈ +4.3% (±33%) | — / — |
| naive-bayes | 25 | ✅ 1024 ms (167 iters) | ✅ 1022 ms (167 iters) | ≈ -0.2% (±32%) | — / — |
| reactors | 21 | ✅ 16285 ms (15 iters) | ✅ 16003 ms (15 iters) | ≈ -1.7% (±5.7%) | — / — |
| reactors | 25 | ✅ 18050 ms (15 iters) | ✅ 18143 ms (15 iters) | ≈ +0.5% (±6%) | — / — |
Internal counter details (ddprof)
ddprof internal counters, latest / dev (✅ = 0, · = unavailable):
| Benchmark | JDK | Dropped rec | Dropped jvmti | Dropped trace | Skipped WC | AGCT fail | Unwind fail |
|---|---|---|---|---|---|---|---|
| akka-uct | 21 | ✅ / ✅ | ✅ / ✅ | ✅ / 5 | 2017 / 2009 | ✅ / ✅ | ✅ / ✅ |
| akka-uct | 25 | ✅ / ✅ | ✅ / ✅ | 1 / 1 | 2344 / 2264 | ✅ / ✅ | ✅ / ✅ |
| finagle-chirper | 21 | ✅ / ✅ | ✅ / ✅ | 3 / 5 | 8749 / 8648 | ✅ / ✅ | ✅ / ✅ |
| finagle-chirper | 25 | ✅ / ✅ | ✅ / ✅ | 3 / 4 | 8234 / 8352 | ✅ / ✅ | ✅ / ✅ |
| fj-kmeans | 21 | ✅ / ✅ | ✅ / ✅ | 2 / 2 | 1239 / 1269 | ✅ / ✅ | ✅ / ✅ |
| fj-kmeans | 25 | ✅ / ✅ | ✅ / ✅ | 2 / 2 | 1267 / 1274 | ✅ / ✅ | ✅ / ✅ |
| future-genetic | 21 | ✅ / ✅ | ✅ / ✅ | 3 / 1 | 3072 / 2901 | ✅ / ✅ | ✅ / ✅ |
| future-genetic | 25 | ✅ / ✅ | ✅ / ✅ | ✅ / 1 | 2974 / 2854 | ✅ / ✅ | ✅ / ✅ |
| naive-bayes | 21 | ✅ / ✅ | ✅ / ✅ | 2 / 2 | 3509 / 3545 | ✅ / ✅ | ✅ / ✅ |
| naive-bayes | 25 | ✅ / ✅ | ✅ / ✅ | 3 / ✅ | 3490 / 3444 | ✅ / ✅ | ✅ / ✅ |
| reactors | 21 | ✅ / ✅ | ✅ / ✅ | 1 / ✅ | 1756 / 1751 | ✅ / ✅ | ✅ / ✅ |
| reactors | 25 | ✅ / ✅ | ✅ / ✅ | ✅ / 1 | 1854 / 1753 | ✅ / ✅ | ✅ / ✅ |
Benchmark Results (commit 277e2d8)Pipeline: https://gitlab.ddbuild.io/DataDog/apm-reliability/benchmarking-platform/-/pipelines/125665685 Commit:
|
| Benchmark | JDK | Latest | Dev | Δ (dev vs latest) | Issues L/D |
|---|---|---|---|---|---|
| akka-uct | 21 | ✅ 10393 ms (21 iters) | ✅ 10229 ms (21 iters) | ≈ -1.6% (±11%) | — / — |
| akka-uct | 25 | ✅ 8931 ms (24 iters) | ✅ 8766 ms (24 iters) | ≈ -1.8% (±9.8%) | — / — |
| finagle-chirper | 21 | ✅ 6009 ms (33 iters) | ✅ 5981 ms (33 iters) | ≈ -0.5% (±24.6%) | |
| finagle-chirper | 25 | ✅ 5507 ms (36 iters) | ✅ 5496 ms (36 iters) | ≈ -0.2% (±24.7%) | |
| fj-kmeans | 21 | ✅ 2761 ms (68 iters) | ✅ 2694 ms (70 iters) | ≈ -2.4% (±2.7%) | — / — |
| fj-kmeans | 25 | ✅ 2757 ms (68 iters) | ✅ 2798 ms (67 iters) | ≈ +1.5% (±2.7%) | — / — |
| future-genetic | 21 | ✅ 2159 ms (86 iters) | ✅ 2039 ms (92 iters) | 🟢 -5.6% | — / — |
| future-genetic | 25 | ✅ 2030 ms (91 iters) | ✅ 2111 ms (88 iters) | 🔴 +4% | — / — |
| naive-bayes | 21 | ✅ 1283 ms (133 iters) | ✅ 1258 ms (136 iters) | ≈ -1.9% (±32.4%) | — / — |
| naive-bayes | 25 | ✅ 1019 ms (168 iters) | ✅ 1020 ms (168 iters) | ≈ +0.1% (±31.8%) | — / — |
| reactors | 21 | ✅ 16429 ms (15 iters) | ✅ 14845 ms (17 iters) | 🟢 -9.6% | — / — |
| reactors | 25 | ✅ 18219 ms (15 iters) | ✅ 18034 ms (15 iters) | ≈ -1% (±4.2%) | — / — |
Internal counter details (ddprof)
ddprof internal counters, latest / dev (✅ = 0, · = unavailable):
| Benchmark | JDK | Dropped rec | Dropped jvmti | Dropped trace | Skipped WC | AGCT fail | Unwind fail |
|---|---|---|---|---|---|---|---|
| akka-uct | 21 | ✅ / ✅ | ✅ / ✅ | 8 / 1 | 1988 / 2062 | ✅ / ✅ | ✅ / ✅ |
| akka-uct | 25 | ✅ / ✅ | ✅ / ✅ | 1 / 9 | 2080 / 2084 | ✅ / ✅ | ✅ / ✅ |
| finagle-chirper | 21 | ✅ / ✅ | ✅ / ✅ | 4 / 3 | 8485 / 8415 | ✅ / ✅ | ✅ / ✅ |
| finagle-chirper | 25 | ✅ / ✅ | ✅ / ✅ | 3 / 2 | 8560 / 8615 | ✅ / ✅ | ✅ / ✅ |
| fj-kmeans | 21 | ✅ / ✅ | ✅ / ✅ | 4 / 2 | 1282 / 1286 | ✅ / ✅ | ✅ / ✅ |
| fj-kmeans | 25 | ✅ / ✅ | ✅ / ✅ | 6 / 1 | 1278 / 1295 | ✅ / ✅ | ✅ / ✅ |
| future-genetic | 21 | ✅ / ✅ | ✅ / ✅ | 2 / 1 | 3002 / 3031 | ✅ / ✅ | ✅ / ✅ |
| future-genetic | 25 | ✅ / ✅ | ✅ / ✅ | 2 / ✅ | 2944 / 2864 | ✅ / ✅ | ✅ / ✅ |
| naive-bayes | 21 | ✅ / ✅ | ✅ / ✅ | 2 / 3 | 3504 / 3524 | ✅ / ✅ | ✅ / ✅ |
| naive-bayes | 25 | ✅ / ✅ | ✅ / ✅ | 2 / 6 | 3485 / 3488 | ✅ / ✅ | ✅ / ✅ |
| reactors | 21 | ✅ / ✅ | ✅ / ✅ | ✅ / ✅ | 1536 / 1691 | ✅ / ✅ | ✅ / ✅ |
| reactors | 25 | ✅ / ✅ | ✅ / ✅ | 1 / ✅ | 1780 / 1837 | ✅ / ✅ | ✅ / ✅ |
Benchmark Results (commit ac8c128)Pipeline: https://gitlab.ddbuild.io/DataDog/apm-reliability/benchmarking-platform/-/pipelines/125736830 Commit:
|
| Benchmark | JDK | Latest | Dev | Δ (dev vs latest) | Issues L/D |
|---|---|---|---|---|---|
| akka-uct | 21 | ✅ 10383 ms (21 iters) | ✅ 10256 ms (21 iters) | ≈ -1.2% (±11.4%) | — / — |
| akka-uct | 25 | ✅ 8923 ms (24 iters) | ✅ 8855 ms (24 iters) | ≈ -0.8% (±10%) | — / — |
| finagle-chirper | 21 | ✅ 5980 ms (33 iters) | ✅ 6028 ms (33 iters) | ≈ +0.8% (±25.1%) | |
| finagle-chirper | 25 | ✅ 5483 ms (36 iters) | ✅ 5498 ms (36 iters) | ≈ +0.3% (±24.1%) | |
| fj-kmeans | 21 | ✅ 2749 ms (68 iters) | ✅ 2717 ms (69 iters) | ≈ -1.2% (±2.7%) | — / — |
| fj-kmeans | 25 | ✅ 2826 ms (66 iters) | ✅ 2762 ms (67 iters) | ≈ -2.3% (±2.6%) | — / — |
| future-genetic | 21 | ✅ 2042 ms (91 iters) | ✅ 2044 ms (90 iters) | ≈ +0.1% (±2.6%) | — / — |
| future-genetic | 25 | ✅ 2130 ms (87 iters) | ✅ 2051 ms (90 iters) | 🟢 -3.7% | — / — |
| naive-bayes | 25 | ✅ 1008 ms (170 iters) | ✅ 1009 ms (169 iters) | ≈ +0.1% (±31.7%) | — / — |
| reactors | 21 | ✅ 16229 ms (15 iters) | ✅ 16543 ms (15 iters) | ≈ +1.9% (±7.5%) | — / — |
| reactors | 25 | ✅ 18332 ms (15 iters) | ✅ 18379 ms (15 iters) | ≈ +0.3% (±3.4%) | — / — |
Internal counter details (ddprof)
ddprof internal counters, latest / dev (✅ = 0, · = unavailable):
| Benchmark | JDK | Dropped rec | Dropped jvmti | Dropped trace | Skipped WC | AGCT fail | Unwind fail |
|---|---|---|---|---|---|---|---|
| akka-uct | 21 | ✅ / ✅ | ✅ / ✅ | 2 / 1 | 2019 / 2112 | ✅ / ✅ | ✅ / ✅ |
| akka-uct | 25 | ✅ / ✅ | ✅ / ✅ | 2 / ✅ | 2418 / 2216 | ✅ / ✅ | ✅ / ✅ |
| finagle-chirper | 21 | ✅ / ✅ | ✅ / ✅ | 5 / 4 | 8420 / 8855 | ✅ / ✅ | ✅ / ✅ |
| finagle-chirper | 25 | ✅ / ✅ | ✅ / ✅ | ✅ / 3 | 8622 / 8796 | ✅ / ✅ | ✅ / ✅ |
| fj-kmeans | 21 | ✅ / ✅ | ✅ / ✅ | 6 / 2 | 1266 / 1280 | ✅ / ✅ | ✅ / ✅ |
| fj-kmeans | 25 | ✅ / ✅ | ✅ / ✅ | 1 / ✅ | 1307 / 1266 | ✅ / ✅ | ✅ / ✅ |
| future-genetic | 21 | ✅ / ✅ | ✅ / ✅ | ✅ / 1 | 3019 / 2888 | ✅ / ✅ | ✅ / ✅ |
| future-genetic | 25 | ✅ / ✅ | ✅ / ✅ | 1 / ✅ | 2890 / 2833 | ✅ / ✅ | ✅ / ✅ |
| naive-bayes | 21 | ✅ / ✅ | ✅ / ✅ | ✅ / ✅ | ✅ / ✅ | ✅ / ✅ | ✅ / ✅ |
| naive-bayes | 25 | ✅ / ✅ | ✅ / ✅ | 4 / 3 | 3505 / 3444 | ✅ / ✅ | ✅ / ✅ |
| reactors | 21 | ✅ / ✅ | ✅ / ✅ | ✅ / ✅ | 1626 / 1497 | ✅ / ✅ | ✅ / ✅ |
| reactors | 25 | ✅ / ✅ | ✅ / ✅ | ✅ / 1 | 1806 / 1830 | ✅ / ✅ | ✅ / ✅ |
There was a problem hiding this comment.
Pull request overview
Introduces a new NativeMem facility to attribute the profiler’s own native memory usage into a small set of categories, tracking per-category live bytes, a moving-window average, and a precise per-category peak; totals are exported via existing counter/JFR pathways.
Changes:
- Added
NativeMemcore implementation (nativeMem.h/.cpp) with per-category live/avg/max tracking and tests. - Instrumented major native allocation sites (calltrace arena/buffers, dictionaries, thread-local data, thread filter, perf mmap, line tables, JFR recording buffers, native symbols gauge) to record alloc/free deltas.
- Emitted totals through the counters table and per-category metrics through
T_DATADOG_COUNTERevents on chunk finish.
Reviewed changes
Copilot reviewed 17 out of 17 changed files in this pull request and generated 2 comments.
Show a summary per file
| File | Description |
|---|---|
| ddprof-lib/src/test/cpp/stringDictionary_ut.cpp | Adds lifecycle accounting-balance test for StringDictionaryBuffer NM_DICTIONARY instrumentation. |
| ddprof-lib/src/test/cpp/nativeMem_ut.cpp | New unit tests covering live/avg/max semantics, peaks between samples, clamping behavior, and category names. |
| ddprof-lib/src/test/cpp/dictionary_ut.cpp | Adds lifecycle accounting-balance test for Dictionary NM_DICTIONARY instrumentation. |
| ddprof-lib/src/main/cpp/threadLocalData.h | Records NM_THREAD_LOCAL on ProfiledThread creation via forTid. |
| ddprof-lib/src/main/cpp/threadLocalData.cpp | Records NM_THREAD_LOCAL decrements on TLS destruction/release paths. |
| ddprof-lib/src/main/cpp/threadFilter.cpp | Accounts for thread-filter chunk and freelist backing allocations under NM_THREAD_FILTER. |
| ddprof-lib/src/main/cpp/stringDictionary.h | Accounts arena chunks + SBTable allocations/frees under NM_DICTIONARY. |
| ddprof-lib/src/main/cpp/profiler.cpp | Accounts per-shard calltrace buffer allocations/frees under NM_CALLTRACE; mirrors native symbol gauge via setLive. |
| ddprof-lib/src/main/cpp/perfEvents_linux.cpp | Accounts perf ring mmap/unmap under NM_PERF. |
| ddprof-lib/src/main/cpp/nativeMem.h | Defines categories and NativeMem API (record/setLive/sample/live/avg/max/totals). |
| ddprof-lib/src/main/cpp/nativeMem.cpp | Implements totals, sampling window averaging, reset, and category naming. |
| ddprof-lib/src/main/cpp/linearAllocator.cpp | Accounts calltrace arena chunk alloc/free under NM_CALLTRACE. |
| ddprof-lib/src/main/cpp/flightRecorder.h | Declares Recording helpers to sample/export native-mem metrics. |
| ddprof-lib/src/main/cpp/flightRecorder.cpp | Samples native-mem each chunk and emits totals + per-category counter events. |
| ddprof-lib/src/main/cpp/dictionary.h | Accounts root table allocation under NM_DICTIONARY. |
| ddprof-lib/src/main/cpp/dictionary.cpp | Accounts key-string alloc/free and overflow/root table alloc/free under NM_DICTIONARY. |
| ddprof-lib/src/main/cpp/counters.h | Adds total native-mem counters (live/avg/max) to the counter table. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| void Recording::updateNativeMemStats() { | ||
| // Refresh the moving-window averages and the observed total peak. Per-category | ||
| // peaks are maintained precisely at allocation time, so they are not sampled | ||
| // here; the total peak is bracketed instead (see writeNativeMem). | ||
| NativeMem::sample(); | ||
|
|
||
| // Mirror the totals into the flat counter table so they flow out through the | ||
| // existing counter path (JFR T_DATADOG_COUNTER events and the JNI debug | ||
| // counters). NATIVE_MEM_MAX_BYTES carries the upper bound on the total peak | ||
| // (sum of precise per-category peaks); the observed lower bound and the | ||
| // per-category values are emitted by writeNativeMem(). | ||
| Counters::set(NATIVE_MEM_LIVE_BYTES, NativeMem::liveTotal()); | ||
| Counters::set(NATIVE_MEM_AVG_BYTES, NativeMem::avgTotal()); | ||
| Counters::set(NATIVE_MEM_MAX_BYTES, NativeMem::maxTotal()); | ||
| } |
| // Per-category live/avg/max, named "<metric>.<category>". The max here is the | ||
| // precise per-category peak tracked at allocation time. | ||
| for (int c = 0; c < NM_NUM_CATEGORIES; c++) { | ||
| NativeMemCategory cat = (NativeMemCategory)c; | ||
| const char *name = NativeMem::categoryName(cat); | ||
| const struct { | ||
| const char *prefix; | ||
| long long value; | ||
| } metrics[] = { | ||
| {"native_mem_live_bytes.", NativeMem::live(cat)}, | ||
| {"native_mem_avg_bytes.", NativeMem::avg(cat)}, | ||
| {"native_mem_max_bytes.", NativeMem::max(cat)}, | ||
| }; |
Benchmark Results (commit 8a9567f)Pipeline: https://gitlab.ddbuild.io/DataDog/apm-reliability/benchmarking-platform/-/pipelines/125902089 Commit: ✅ Within expected boundariesNo significant runtime deltas (all within run-to-run noise) and no internal-counter outliers. Runtime details (per benchmark × JDK)
Internal counter details (ddprof)ddprof internal counters, latest / dev (✅ = 0, · = unavailable):
|
Add a NativeMem facility that tracks the profiler's own native memory
usage per category, with a moving-window average and a running peak.
- NativeMemCategory enum whose per-category live gauges partition the
total (each backing allocation belongs to exactly one category, so
there is no double counting).
- Live accounting is always-on and independent of the COUNTERS build
flag: record() is a single relaxed atomic add, async-signal-safe and
usable from signal handlers.
- sample() folds the live gauges into a moving-window average and a
high-water max; it is ticked once per JFR chunk finish.
- Totals mirror into the existing NATIVE_MEM_{LIVE,AVG,MAX}_BYTES
counters (JFR + JNI counter path); per-category values are emitted as
native_mem_{live,avg,max}_bytes.<category> counter events, reusing the
existing counter event format (no new event type).
Instrumented sites (tagged CALLTRACE, their sole use today): the
LinearAllocator chunk alloc/free (the call-trace arena) and the
per-shard calltrace buffers. Accounting lives with the semantic owner
rather than OS::safeAlloc, which stays category-agnostic.
First-cut limits, documented in code: only CALLTRACE sites are
instrumented so far (other categories read 0 until tagged); the max is
sampled rather than spike-accurate; transient negative live is clamped
to 0. The existing reserved/used/waste counters (CALLTRACE_STORAGE_BYTES,
DICTIONARY_ARENA_WASTE_BYTES) remain an independent nested dimension and
are intentionally not summed into the per-category total.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Track each category's peak at allocation time instead of sampling it at chunk finish, so a spike that rises and falls between two ticks is still captured. record() updates the per-category high-water mark on positive deltas via a relaxed CAS: the common (no new peak) path is a single load plus a compare, and the CAS fires only when a genuinely higher peak is set, which is rare since the peak is monotonic. Frees skip the check. The total peak avoids a shared global counter (a contention hotspot on a single cache line hit by every allocation) and is instead reported as a bracket: - NATIVE_MEM_MAX_BYTES = upper bound = sum of the precise per-category peaks (exact when the peaks coincide, otherwise an overestimate). - native_mem_max_observed_total_bytes = lower bound = the largest instantaneous total seen at a sampling tick. sample() no longer touches the per-category peaks; it only refreshes the moving averages and the observed total. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Extend per-category accounting beyond CALLTRACE to the other large, cleanly-paired backing allocations: - THREAD_LOCAL: `new`/`delete ProfiledThread` (sizeof, per thread). - JFR_BUFFERS: `new`/`delete Recording` (embeds the RecordingBuffer array and cpu-monitor buffer). - LINE_TABLES: the malloc'd JVMTI line-number table copy, tracked beside the existing LINE_NUMBER_TABLES counter and freed in ~SharedLineNumberTable (byte size recovered from the stored entry count). - PERF: the perf ring mmap (2 * page_size), paired with its munmap. - THREAD_FILTER: ChunkStorage chunks (bounded, tagged only on successful CAS install) and the FreeListNode array. - CODECACHE: a recomputed gauge, mirrored via the new NativeMem::setLive() at the existing CodeCache size-set site. setLive() overwrites live and still advances the peak. Not tagged in this commit: - CONTEXT has no separate allocation — the OTel context record is embedded in ProfiledThread, so it is already counted under THREAD_LOCAL; tagging it again would double-count. The category stays 0 by design. - DICTIONARY is deferred to its own commit: it has two implementations (the older per-key-malloc Dictionary used for symbols/packages, with set()-based counters and bulk key-frees that lose per-item size, and the arena-based StringDictionary where keys live inside chunks). A partial number that tagged only one would mislead, so it needs dedicated per-implementation handling. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add DICTIONARY to the per-category accounting, and remove the CONTEXT category. CONTEXT had no separate backing allocation — the OTel context record is embedded in ProfiledThread and already counted under THREAD_LOCAL — so it would have been a permanently-zero line. Removed rather than left dangling. DICTIONARY is tagged across both implementations, counting physical backing allocations once each (single category; per-role breakdown, if ever wanted, belongs in a nested dimension, not extra top-level categories): - Dictionary (older, per-key malloc; used for symbols/packages): root and overflow DictTables (constant size) and key strings. Key frees are bulk and lose per-item sizes, but keys are null-terminated so the malloc'd size (strlen + 1) is recovered at free — no running total needed. - StringDictionary / StringArena (arena-based): arena chunks and root / overflow SBTables. Keys are bump-allocated inside chunks, so they are NOT counted separately (that would double-count). Chunk and SBTable accounting is unconditional, independent of the diagnostic counters' _counter_offset gate, so anonymous dictionaries are covered too. Tests: lifecycle invariants for both implementations assert the accounting grows on insert, returns exactly to the construction baseline after clear(), and to zero after destruction — directly proving inc/dec pairing (including the strlen-at-free path and arena chunk growth). Full gtestDebug suite green (356 tests). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The class comment implied every record() call needs to be async-signal-safe. In fact most categories allocate via malloc/new off the signal path, where the property is irrelevant. Only the CALLTRACE arena allocates from within the sampling signal handler (via OS::safeAlloc's raw mmap syscall), so that is where record() staying a relaxed atomic add actually matters. Reword accordingly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The "CodeCache" name is async-profiler's, and collides with the JVM's JIT code cache. This category measures something unrelated: the profiler's own per-native-library symbol tables (used to symbolicate native frames), not JVM-managed code. Rename the NativeMem category and its JFR label to native_symbols so the metric is unambiguous. The existing CODECACHE_NATIVE_SIZE_BYTES counter keeps its name for continuity; only the new NativeMem category is renamed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
8a9567f to
3c82245
Compare
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 17 out of 17 changed files in this pull request and generated 2 comments.
Comments suppressed due to low confidence (1)
ddprof-lib/src/main/cpp/flightRecorder.cpp:1815
- writeNativeMem() serializes values with Buffer::putVar64(u64). If any native-mem gauge goes negative (possible until all free sites are instrumented), the implicit signed->unsigned conversion will emit a huge varint. Clamp to 0 before encoding to keep the JFR/counter stream valid.
auto emit = [&](const char *label, long long value) {
int start = buf->skip(1);
buf->putVar64(T_DATADOG_COUNTER);
buf->putVar64(_start_ticks);
buf->putUtf8(label);
buf->putVar64(value);
writeEventSizePrefix(buf, start);
flushIfNeeded(buf);
| long long NativeMem::liveTotal() { | ||
| long long total = 0; | ||
| for (int c = 0; c < NM_NUM_CATEGORIES; c++) { | ||
| total += load(_live[c]); | ||
| } | ||
| return total; | ||
| } |
| @@ -1383,13 +1385,17 @@ Error Profiler::start(Arguments &args, bool reset) { | |||
| return Error("Not enough memory to allocate stack trace buffers (try " | |||
| "smaller jstackdepth)"); | |||
| } | |||
Benchmark Results (commit 3c82245)Pipeline: https://gitlab.ddbuild.io/DataDog/apm-reliability/benchmarking-platform/-/pipelines/125929510 Commit: ✅ Within expected boundariesNo significant runtime deltas (all within run-to-run noise) and no internal-counter outliers. Runtime details (per benchmark × JDK)
Internal counter details (ddprof)ddprof internal counters, latest / dev (✅ = 0, · = unavailable):
|
The class comment still described record() as "a single relaxed atomic add" from before precise per-category max was added. record() now also does a conditional lock-free high-water update on allocation. Correct the wording; the async-signal-safety guarantee still holds (lock-free atomics). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 17 out of 17 changed files in this pull request and generated no new comments.
Comments suppressed due to low confidence (1)
ddprof-lib/src/main/cpp/flightRecorder.cpp:1801
- NativeMem::sample() explicitly clamps negative per-category live values to 0 when computing totals/averages, but the exported total live counter uses NativeMem::liveTotal(), which can go negative if any category is temporarily negative (the scenario sample() is already designed to tolerate). This can produce a nonsensical negative "native_mem_live_bytes" in the counter stream and makes the totals inconsistent with the sampled/clamped window.
// per-category values are emitted by writeNativeMem().
Counters::set(NATIVE_MEM_LIVE_BYTES, NativeMem::liveTotal());
Counters::set(NATIVE_MEM_AVG_BYTES, NativeMem::avgTotal());
Counters::set(NATIVE_MEM_MAX_BYTES, NativeMem::maxTotal());
Two invariants the accounting relies on are now asserted (stripped under NDEBUG, so no release cost and never in a real signal handler): - Per-category live bytes never go negative: record() asserts the post-update value >= 0. A negative means an unbalanced/oversized free. This lets liveTotal() sum without clamping, since each term is >= 0. - Dictionary keys are NUL-free strings of exactly `length`: allocateKey() asserts strlen == length. The NM_DICTIONARY free path recovers a key's size via strlen at clear(), so an embedded NUL would under-count; the assert trips in debug/gtest instead of silently drifting. The sample() clamp is retained as a release-mode safety net for the asserted-impossible negative case, and its comment updated to say so. The old NegativeLiveClampedInSample test is removed: it deliberately drove a category negative, which now (correctly) trips the record() assert. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 19db9d3172
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| if (rec != nullptr) { | ||
| // NULL first, deallocate later | ||
| _rec = nullptr; | ||
| NativeMem::record(NM_JFR_BUFFERS, -(long long)sizeof(Recording)); |
There was a problem hiding this comment.
Keep JFR buffer bytes counted through final chunk
When stopping a recording, delete rec invokes Recording::~Recording(), which calls finishChunk(true) and emits the native-memory counters. This decrement runs before that destructor, so the final JFR chunk reports jfr_buffers and the total as if the Recording buffers were already freed even though they are still live during serialization; move the decrement after delete rec or after finishChunk to keep stop output accurate.
Useful? React with 👍 / 👎.
| // CodeCache here is the profiler's native-symbol tables (not the JVM code | ||
| // cache). Its size is a recomputed gauge (not alloc/free deltas), so mirror | ||
| // it as an absolute. It shares the accuracy caveats of the counter above. | ||
| NativeMem::setLive(NM_NATIVE_SYMBOLS, native_libs.memoryUsage()); |
There was a problem hiding this comment.
Refresh native-symbol gauge on stop as well
When users stop without a prior dump, execution follows ACTION_STOP into Profiler::stop, which refreshes Libraries and then calls _jfr.stop() without executing this dump-only NativeMem::setLive. As a result the final recording emits native_mem_*_bytes.native_symbols as 0 or stale for the normal stop path; mirror native_libs.memoryUsage() before _jfr.stop() too, or centralize this refresh in the chunk-finish path.
Useful? React with 👍 / 👎.
| Counters::set(CODECACHE_NATIVE_SIZE_BYTES, native_libs.memoryUsage()); | ||
| Counters::set(CODECACHE_RUNTIME_STUBS_SIZE_BYTES, | ||
| native_libs.memoryUsage()); | ||
| // CodeCache here is the profiler's native-symbol tables (not the JVM code | ||
| // cache). Its size is a recomputed gauge (not alloc/free deltas), so mirror | ||
| // it as an absolute. It shares the accuracy caveats of the counter above. | ||
| NativeMem::setLive(NM_NATIVE_SYMBOLS, native_libs.memoryUsage()); |
|
|
||
| // Every category exposes a distinct, non-empty name. |
Benchmark Results (commit 19db9d3)Pipeline: https://gitlab.ddbuild.io/DataDog/apm-reliability/benchmarking-platform/-/pipelines/125953166 Commit: ✅ Within expected boundariesNo significant runtime deltas (all within run-to-run noise) and no internal-counter outliers. Runtime details (per benchmark × JDK)
Internal counter details (ddprof)ddprof internal counters, latest / dev (✅ = 0, · = unavailable):
|
What
Adds a
NativeMemfacility to measure the profiler's own native memory usage — currently we only observe whole-process RSS. Per-category live gauge, moving-window average, and a precise per-category peak, wired through the existing counter/JFR path.Why
We measure the agent's native footprint only as an RSS delta (~150–200 MB) with no breakdown. This gives an in-process, attributable, categorized number we can trend, and a foundation for later pinpointing which code is responsible.
How
NativeMem(nativeMem.h/nativeMem.cpp): aNativeMemCategoryenum whose per-category live gauges partition the total — each backing allocation belongs to exactly one category, so no double-counting.COUNTERSbuild flag:record()is a single relaxed atomic add. Most sites are off the signal path (malloc/new is not async-signal-safe anyway); the exception is CALLTRACE, whose arena allocates inside the sampling signal handler (viaOS::safeAlloc), sorecord()is kept async-signal-safe there.record()maintains each category's high-water mark at allocation time via a relaxed CAS (common path is load+compare; CAS only on a new peak; frees skip it). Spikes between sample ticks are still captured.sample()refreshes the moving-window averages and the observed total, ticked once per JFR chunk finish.NATIVE_MEM_{LIVE,AVG,MAX}_BYTES(JFRT_DATADOG_COUNTER+ JNI debug counters); per-categorynative_mem_{live,avg,max}_bytes.<category>values reuse the same event format — no new event type.Total peak: bracketed, no shared hotspot
A precise total peak would need a single global atomic hit by every allocation (a contention hotspot). Instead the total peak is bracketed:
NATIVE_MEM_MAX_BYTES= upper bound = sum of the precise per-category peaks.native_mem_max_observed_total_bytes= lower bound = the largest instantaneous total seen at a sampling tick.Instrumented categories
Tagged via
record()at their backing alloc/free sites (precise, always-on):LinearAllocatorchunks (call-trace arena) + per-shard calltrace buffers.Dictionary(symbols/packages — root/overflowDictTables + key strings, whose size is recovered viastrlenat bulk-free), and the arenaStringDictionary(arena chunks + root/overflowSBTables; keys live inside chunks so they're not counted separately). Counted unconditionally, independent of the diagnostic counters' offset gate. Single category — a per-role split, if ever wanted, belongs in a nested dimension.ProfiledThreadper thread.Recordingobject (embeds the JFR buffer array).mmap(2 × page_size).ChunkStoragechunks (bounded) + free-list array.Gauge-mirrored (
NativeMem::setLive()at a recompute site):CodeCache; not the JVM code cache). Shares the existingCODECACHE_NATIVE_SIZE_BYTEScounter's accuracy caveats (under-counts symbol-name strings, DWARF tables, and the blob array; stale snapshot); those are fixed in fix(codecache): make memoryUsage() accurate and live #677, which this gauge picks up automatically once both land.Not yet covered (reads 0):
new/mallocacross the remaining files → best captured by malloc interception, not hand-tagging.(
CONTEXTwas removed: the OTel context record is embedded inProfiledThread, already counted underTHREAD_LOCAL; a separate category would double-count or read a permanent 0.)So the total is still a subset of the RSS delta until interception lands, but now covers the major consumers.
Avoiding double-counting
Call-trace bytes appear at three layers (
safeAllocmmap,LINEAR_ALLOCATOR_BYTESreserved chunks,CALLTRACE_STORAGE_BYTESused-within-chunks); summing counters would double/triple-count.NativeMemcounts backing memory once per category; the existing reserved/used/waste counters remain an independent nested diagnostic dimension. The dictionary tagging follows the same rule — e.g. arena keys are inside chunks and are not counted separately.Performance
Designed to be off the hot path. Measured in operations, not observed as a regression in this workload:
CallTraceStorage::putnormally just bump-allocates within an existing arena chunk and touches norecord().record()runs only when the call-trace arena grows a new chunk (8 MiB apart), and there it's a lock-free relaxed atomic add + high-water update — async-signal-safe.record()is one relaxedfetch_add; on allocation it also does a relaxed load + compare, with a CAS only when a new per-category peak is set (rare, since the peak is monotonic). This sits next to amalloc/calloc/mmap/newthat dominates the cost, so it's negligible. These sites (dictionary inserts, chunk/table growth, thread/lib/recording creation) are not per-sample.sample()is O(categories × window) ≈ 9 × 64 folds;writeNativeMem()emits ~28 counter events. This runs once per chunk rotation (seconds–minutes), off the sampling path.clear(): the per-key free now also does astrlento recover the freed size, soclear()cost is O(sum of key lengths) rather than O(key count).clear()already walks every key; this is a marginal constant-factor increase, at dictionary rotation only.No new locks, no shared global counter on the hot path (the total peak is bracketed precisely to avoid one), and the always-on atomics are relaxed.
sample()/writeNativeMem()run within the existing recording flush.Roadmap
Testing
nativeMem_ut(8 tests): live tracking + total partition, single-sample avg==live, moving average, max high-water, precise per-category max catching an inter-sample spike (with the bracket), negative-live clamping,setLivegauge + peak retention, category names.dictionary_ut,stringDictionary_ut): accounting grows on insert, returns exactly to the construction baseline afterclear(), and to zero after destruction — directly proving inc/dec pairing (incl. the strlen-at-free path and arena chunk growth).:ddprof-lib:gtestDebugsuite green — 359 tests, 0 failures (rebased onto latestmain).🤖 Generated with Claude Code