---
title: "Spatially Stratified Estimation with Sections"
output:
  rmarkdown::html_vignette:
    highlight: null
vignette: >
  %\VignetteIndexEntry{Spatially Stratified Estimation with Sections}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include = FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>"
)
```

## Introduction

A spatially stratified creel survey divides the lake into discrete geographic
sections — for example, North, Central, and South — each managed as a separate
reporting unit. Each section has its own observed count data and its own set of
angler interviews, so effort levels and catch rates can differ materially between
sections.

`add_sections()` is the right tool when the survey design *explicitly* stratifies
by section: when different field crews patrol different areas, when section-level
estimates are required in the final report, or when the lake is large enough that
assuming uniform effort across the full water body would be misleading. If you are
simply curious about spatial patterns but did not design the survey with sections
in mind, use the standard `creel_design()` workflow without `add_sections()`.

---

## Building a Sectioned Design

```{r design, warning = FALSE}
library(tidycreel)

data(example_sections_calendar)
data(example_sections_counts)
data(example_sections_interviews)

sections_df <- data.frame(
  section = c("North", "Central", "South"),
  stringsAsFactors = FALSE
)

design <- creel_design(example_sections_calendar, date = date, strata = day_type)
design <- add_sections(design, sections_df, section_col = section)
design <- add_counts(design, example_sections_counts)
design <- add_interviews(
  design, example_sections_interviews,
  catch = catch_total,
  effort = hours_fished,
  harvest = catch_kept,
  trip_status = trip_status,
  trip_duration = trip_duration
)
print(design)
```

The design now shows three registered sections (North, Central, South). All
downstream estimators automatically detect sections and return per-section output.

---

## Per-Section Effort Estimation

`estimate_effort()` detects the registered sections and returns one row per
section plus a `.lake_total` row that aggregates across sections. The lake-total
standard error accounts for cross-section covariance (described in the Variance
Aggregation section below).

```{r effort}
effort_est <- estimate_effort(design)
print(effort_est$estimates)
```

Central has the highest fishing effort among the three sections, while South has
the lowest. The `prop_of_lake_total` column shows each section's share of total
angler-hours for the survey period, and `se_prop_of_lake_total` its standard
error, so shares can be compared with the same regard for precision as the
totals themselves. Both come from one `survey::svyratio()` call, which accounts
for the correlation between a section's total and the lake-wide total that
contains it. On the `.lake_total` row the share is exactly 1 with a standard
error of 0 — that row's share of itself is structural, not estimated.

---

## Per-Section Catch Rate

Catch rate (CPUE, fish per angler-hour) is a ratio estimator. Ratios are not
additive across sections: you cannot average North's rate and South's rate to
produce a valid lake-wide CPUE. For this reason, `estimate_catch_rate()` on a
sectioned design returns one row per section with **no `.lake_total` row**.

```{r catch_rate, warning = FALSE}
cpue_est <- estimate_catch_rate(design)
print(cpue_est$estimates)
```

::: {.alert .alert-warning}
**Notice there is no `.lake_total` row.** Catch rate is a ratio estimator and
ratios are not additive — South's high catch rate cannot simply be averaged
with North's low rate to produce a valid lake-wide CPUE. To estimate the
lake-wide catch rate, call `estimate_catch_rate()` on a design without section
registration.
:::

```{r catch_rate_lake, warning = FALSE}
# Lake-wide catch rate: build an unsectioned design using the same data
design_nosections <- creel_design(example_sections_calendar, date = date, strata = day_type)
design_nosections <- add_counts(design_nosections, example_sections_counts)
design_nosections <- add_interviews(
  design_nosections, example_sections_interviews,
  catch = catch_total, effort = hours_fished,
  harvest = catch_kept, trip_status = trip_status, trip_duration = trip_duration
)
estimate_catch_rate(design_nosections)$estimates
```

---

## Per-Section and Lake-Wide Total Catch

Total catch combines effort and catch rate within each section:
`TC_i = E_i × CPUE_i`. The lake-wide total is `sum(TC_i)` — **not**
`E_total × CPUE_pooled`. With `aggregate_sections = TRUE` (the default), a
`.lake_total` row is appended to the output.

```{r total_catch, warning = FALSE}
catch_est <- estimate_total_catch(design, aggregate_sections = TRUE)
print(catch_est$estimates)
```

South has a high catch rate but the lowest effort, which limits its contribution
to the lake total. Central has a moderate catch rate combined with the highest
effort, making it the largest contributor. The `prop_of_lake_total` column shows
each section's share of the lake-wide total catch, and `se_prop_of_lake_total`
its standard error, so the shares can be compared against their uncertainty
rather than read as exact. The error is derived from the same section variances
the `.lake_total` row's own standard error is built from; the `.lake_total` row
reports `0`, because its share of itself is 1 by construction rather than
estimated.

```{r total_harvest, warning = FALSE}
harvest_est <- estimate_total_harvest(design, aggregate_sections = TRUE)
print(harvest_est$estimates)
```

`example_sections_interviews` carries no party-size column, so both calls above
raise a warning that the rate they multiply is per *party*-hour while the
count-derived effort is angler-hours. It is correct for this dataset and is
hidden here only by the chunks' `warning = FALSE`, which predates it. With one
angler per party the totals are right as printed; with larger parties they are
understated. Supply `n_anglers` from your own interview data, or state a
constant party size (`n_anglers = 1`) when every interview really is a single
angler, to remove the ambiguity.


---

## Variance Aggregation — Why `method = "correlated"` Is the Default

In a standard shared-calendar creel design, the field crew works the entire lake
on each survey day. Every section is counted and interviewed on the same calendar
days: if a day is sampled, all three sections are observed; if a day is not
sampled, none are.

Sharing survey days creates cross-section covariance in the sampling errors. In
practice this covariance tends to be negative: on high-effort days all sections
run high together, and on low-effort days they run low together. The
shared-calendar variance formula accounts for this covariance and produces a
narrower lake-total standard error than simply adding the section standard errors
in quadrature.

`method = "independent"` is appropriate only when sections are surveyed by
separate, uncoordinated crews — for example, when the South section is run by
Crew A on different days from the North section (Crew B). In that case the
sections have truly independent sampling errors and Cochran (1977, §5.2)
additivity applies: `SE_total = sqrt(sum(SE_h^2))`.

**Rule of thumb:** If your crew drives the entire lake on each survey day, use
the default `method = "correlated"`. If different sections have separate,
non-overlapping calendars, use `method = "independent"`.

```{r variance_methods}
# Default: accounts for cross-section covariance (shared-calendar designs)
estimate_effort(design, method = "correlated")$estimates

# Alternative: Cochran 5.2 additivity (independent crews, non-overlapping calendars)
estimate_effort(design, method = "independent")$estimates
```

---

## Missing Section Warning

If a registered section has no count data on any survey day, `estimate_effort()`
produces an NA row with `data_available = FALSE` and emits a warning. This
prevents silent omission of sections that should have been observed.

```{r missing_section, warning = TRUE}
north_central_counts <- example_sections_counts[
  example_sections_counts$section != "South",
]

design_missing <- creel_design(example_sections_calendar, date = date, strata = day_type)
design_missing <- add_sections(design_missing, sections_df, section_col = section)
design_missing <- suppressWarnings(add_counts(design_missing, north_central_counts))

estimate_effort(design_missing, missing_sections = "warn")$estimates
```

To turn warnings into errors (for automated pipelines), use
`missing_sections = "error"`.
