rtransparency 1.2.0

A correctness, speed and usability release. Detection changes were measured article by article on every labeled set (data-raw/benchmark/snapshot.R and evaluate.R); behavior-preserving changes were verified to give identical predictions on 3124 cached articles.

Bug fixes

Measured effect

rt_all_pmc() before (1.1.0) and after, on the labeled sets (sensitivity / specificity, %):

Indicator Labeled set 1.1.0 1.2.0
Replication enriched sample (111 positives) 92.8 / 34.5 96.4 / 33.1
Replication 2023 sample (17 positives) 82.4 / 98.5 82.4 / 98.4
Data sharing 2023 sample (123 positives) 90.2 / 98.1 91.9 / 98.1
Code sharing held-out set (109 positives) 86.2 / 98.6 88.1 / 99.5
Code sharing 2023 sample (33 positives) 93.9 / 99.1 93.9 / 99.0
Conflicts of interest 2023 sample (927 positives) 100 / 91.8 100 / 90.4

Funding, registration, novelty, open access, AI disclosure and reporting guidelines are unchanged on every labeled set, as are conflicts of interest on the independently labeled held-out set. The code row moves because rt_all_pmc() now matches the standalone detectors, which the published benchmarks already used. Reporting stays at 95.4 / 99.0 for rt_all_pmc(), and rt_reporting_pmc() rises from 93.8 to 95.4 because citation markers no longer hide guideline names.

The one conflict-of-interest change in the 2023 sample is an article whose footnote reads “Declaration of competing interest: None.” but whose label is FALSE. Those labels were reconciled against the 1.1.0 detector’s output, so they inherit its misses; three more articles with a “Competing interests: …” statement labeled FALSE explain most of the plain-text specificity in results_txt_parity.md. The labels are left as they are, pending the maintainer’s review, and the new blind 2025 rounds are the fix. Outside the labeled sets, the AI fixes catch two more genuine disclosures, the data vetoes remove four false positives, and French conflict-of-interest detection on the multilingual corpus rises from 30% to 77%.

Regenerating every benchmark report for this release also exposed drift that predates it: on the held-out Serghiou et al. (2021) set, the 1.1.0 detectors (like this release’s) score funding at 91.7% sensitivity and registration at 92.7% specificity, where the report last generated for 0.9.0 said 100% and 96.9%. The reports and rt_accuracy now show the current values; finding the changes responsible is on the roadmap.

Performance

New features

Methodology

Other changes

rtransparency 1.1.0

Not released to CRAN; these changes ship in 1.2.0.

Two new transparency indicators, bringing the total to ten.

rtransparency 1.0.0

First stable release, and a rename.

rtransparent 0.9.11

Citation, documentation, and packaging polish.

rtransparent 0.9.10

Replication is now accuracy-corrected; a fresh validation of replication and AI.

rtransparent 0.9.9

Conflict-of-interest and funding detection in five more languages.

rtransparent 0.9.8

The plain-text detectors now share the PMC detection logic.

rtransparent 0.9.7

Corpus-scale batch processing.

rtransparent 0.9.6

The hand-labeled 2023 validation sample reaches 1000 articles.

rtransparent 0.9.5

The hand-labeled 2023 validation sample is expanded to 980 articles (265 new), with a focused improvement to replication precision and a further funding fix.

rtransparent 0.9.4

The hand-labeled 2023 validation sample is expanded to 715 articles (210 new), with three small detector fixes surfaced by the new batches.

rtransparent 0.9.3

The hand-labeled 2023 validation sample is expanded to 505 articles (120 new), with three small detector fixes surfaced by the new batches.

rtransparent 0.9.2

A precision release from the next round of hand-label review (2023 sample grown to 385 articles).

rtransparent 0.9.1

This release overhauls the novelty detector for both recall and precision, fixes two long-standing bugs in the public PMC entry points, and corrects mislabeled articles in the 2023 validation sample.

rtransparent 0.9.0

This is a feature release centered on the novelty and replication detectors and a second, independent validation set.

rtransparent 0.8.16

rtransparent 0.8.15

rtransparent 0.8.14

rtransparent 0.8.13

rtransparent 0.8.12

rtransparent 0.8.11

rtransparent 0.8.10

rtransparent 0.8.9

rtransparent 0.8.8

rtransparent 0.8.7

Fixes for genome data-papers (Darwin Tree of Life and similar), found during the manual validation of 1,000 open-access PMC articles:

rtransparent 0.8.6

rtransparent 0.8.5

rtransparent 0.8.4

Documentation and example data, so the package website showcases every indicator:

rtransparent 0.8.3

Further fixes from the manual validation on a fresh sample of 1,000 open-access PMC articles from 2023:

rtransparent 0.8.2

rtransparent 0.8.1

Fixes from a manual validation on a fresh, disjoint sample of 1,000 open-access PMC articles from 2023:

rtransparent 0.8.0

rtransparent 0.7.1

rtransparent 0.7.0

Improvements from a large audit: the tool was run over 1,000 cached open-access PMC articles and a sample was hand-checked against the human-labeled benchmark.

rtransparent 0.6.1

Precision and recall fixes from an independent manual review of a sample of open-access PMC articles:

rtransparent 0.6.0

rtransparent 0.5.1

rtransparent 0.5.0

rtransparent 0.4.3

rtransparent 0.4.2

rtransparent 0.4.1

rtransparent 0.4.0

rtransparent 0.3.4

rtransparent 0.3.3

rtransparent 0.3.2

rtransparent 0.3.1

rtransparent 0.3.0

rtransparent 0.2.5

rtransparent 0.2

rtransparent 0.1

mirror server hosted at Truenetwork, Russian Federation.