---
title: "Getting started with trialdiff"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Getting started with trialdiff}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include = FALSE}
knitr::opts_chunk$set(collapse = TRUE, comment = "#>")
```

## Why trialdiff?

Clinical trial datasets are regenerated many times during a study. Between two
data cuts, observations are added, removed or corrected, derivations change, and
metadata evolves. Generic tools such as `diffdf` and `waldo` can tell you *that*
two datasets differ. `trialdiff` answers the clinical-programming question that
comes next:

> What changed, what does the change represent, and which downstream analyses or
> outputs might be affected?

`trialdiff` is organised into five layers, each usable on its own:

1. **Compare** - [compare_cut()]
2. **Classify** - [classify_changes()]
3. **Trace** - [define_lineage()], [trace_dependencies()]
4. **Assess** - [assess_impact()]
5. **Report** - [report_diff()]

All five layers are deterministic and rule-based. Nothing is inferred by an
opaque model, and no statistical significance is ever claimed.

## The example data

The package ships synthetic `ADSL` and `ADLB` data cuts. No proprietary data is
used.

```{r}
library(trialdiff)
str(adsl_cut1, max.level = 1)
```

## 1. Compare two data cuts

```{r}
diff <- compare_cut(
  old = adsl_cut1,
  new = adsl_cut2,
  by = "USUBJID",
  dataset = "ADSL"
)
diff
```

The result is a `tdiff` object with `added`, `removed`, `modified` and `schema`
tables. For example, the treatment-assignment change:

```{r}
diff$modified[, c("USUBJID", "variable", "old_value", "new_value", "change")]
```

## 2. Classify the changes

```{r}
classified <- classify_changes(diff)
table(classified$register$category_label)
```

Every classification carries a plain-language explanation:

```{r}
classified$register$reason[classified$register$category == "treatment_assignment_change"]
```

## 3. Define lineage

Lineage is explicit metadata. Nodes are datasets (`ADSL`), variables
(`ADSL.TRT01P`), analyses (`MMRM`) or outputs (`Table_14_2_1`).

```{r}
lineage <- define_lineage(
  lineage_edge("ADSL.TRT01P", "ADLB.TRT01P", relationship = "groups_by"),
  lineage_edge("ADLB.AVAL", "MMRM", relationship = "models"),
  lineage_edge("MMRM", "Table_14_2_1", relationship = "reports")
)
trace_dependencies(lineage, from = "ADSL.TRT01P")
```

## 4. Assess impact

```{r}
impact <- assess_impact(classified, adsl_adlb_lineage)
impact$impacts[, c("node", "node_type", "level", "requires_rerun")]
```

Impact is graded as `definitely_affected`, `potentially_affected` or
`unlikely`. Analyses and outputs are flagged as requiring review/rerun; the
package never claims statistical impact.

## 5. Report

```{r}
report <- report_diff(diff, impact = impact, output = "list")
report$data$review_items
```

Use `output = "report.html"` to write a self-contained HTML report, or
`as_json()` for machine-readable output suitable for automated QC pipelines.

## Next steps

* `vignette("change-classification")` for the rule system.
* `vignette("lineage-and-impact")` for the lineage model.
* `vignette("ecosystem")` for how `trialdiff` complements existing tools.
