---
title: "Progressive Count Surveys"
output:
  rmarkdown::html_vignette:
    highlight: null
vignette: >
  %\VignetteIndexEntry{Progressive Count Surveys}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include = FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>"
)
```

## Overview

A progressive count survey replaces the random spot-check of an instantaneous count with
a complete traversal of the access corridor. The observer drives or walks the entire
shoreline, access road, or boat ramp circuit — counting every angler encountered along the
way. This vignette demonstrates the full progressive count workflow: scheduling,
data collection structure, effort estimation, and catch estimation.

## Instantaneous vs. Progressive Counts

| Feature | Instantaneous | Progressive |
|---------|---------------|-------------|
| Observer visits | One spot at a random time | Full circuit of the section |
| Count meaning | Anglers visible at a moment | Anglers encountered during circuit |
| Estimator | $\hat{E}_d = C \times T_d$ | $\hat{E}_d = C \times T_d$ |
| Additional input | — | `circuit_time` (τ) and `period_length_col` ($T_d$) |
| Circuit time τ | Not needed | Required (cancels algebraically) |

The estimators are algebraically identical once the progressive count formula is
expanded: $\hat{E}_d = C \times \tau \times (T_d / \tau) = C \times T_d$.
The circuit time $\tau$ cancels, so only the raw count $C$ and the total open hours
$T_d$ enter the final calculation. Nevertheless, $\tau$ must be supplied to
`add_counts()` as a check that the field protocol (circuit duration) is documented.

## When to Use Progressive Counts

Progressive counts are preferred when:

- Anglers are spread along a linear corridor (river, canal, reservoir shoreline road)
  and a single spot-check would miss substantial activity
- The access route is fully enumerable in one traverse within a count period
- The observer can travel the circuit faster than anglers move between fishing spots
  (see the conditions below — this is a requirement of the method, not a preference)

Instantaneous counts are preferred when:

- The waterbody is large and a complete traverse is impractical
- Multiple sections are counted simultaneously by separate crews
- Observer presence along the circuit might disturb anglers or alert them to upcoming
  interviews

## Conditions for an Unbiased Estimate

$\hat{E}_d = C \times T_d$ is unbiased only when the field protocol satisfies three
conditions (Hoenig et al. 1993, p. 725). They are properties of how the count is
*collected*, so no check inside `add_counts()` can confirm them — they are the
analyst's responsibility:

1. **The starting location on the circuit is chosen randomly.** Always beginning at
   the boat ramp biases the count toward whatever is happening there at that hour.
2. **The direction of travel is chosen randomly** from the two alternatives.
   `generate_progressive_start()` draws this for you and returns it in the
   `direction` column — it is a field instruction, not decoration.
3. **The observer travels faster than the anglers move** while those anglers are
   fishing. If anglers can outpace the observer they may be counted twice or missed.

Two further cautions from the same paper:

- **Do not interrupt the circuit to conduct interviews.** The derivation assumes
  constant travel speed; stopping to interview breaks it and is one of the four
  errors Hoenig et al. single out. Interview on a separate pass.
- **The count estimates angler-days of effort, not the number of trips.** Reading
  $C$ as a trip count "results in a negative bias that can be severe"
  (Hoenig et al. 1993). $C$ counts anglers present, and an angler fishing all day
  is counted once per circuit — not once per trip.

Scheduling the circuit start correctly is the fourth condition, and it is the subject
of the section below.

## Data Requirements

A progressive count dataset needs:

1. **Calendar** — sampled dates with `day_type`. The calendar carries no hours
   column: `creel_design()` reads only the date and the strata, so $T_d$ must
   travel with the counts.
2. **Count data** — one row per sampled day with raw angler count ($C$) and
   a column recording $T_d$, passed as `period_length_col`. This is the only
   place the estimator looks for the open hours.
3. **`circuit_time`** — a single numeric value (hours) for the traversal duration $\tau$

## A Complete Example

### Survey Setup

We survey a 25-km reservoir access road over a 4-week season (July–July). A single
crew completes the full circuit in 2 hours ($\tau = 2$ h). The access road is open
10 hours per day.

```{r setup-design, message = FALSE}
library(tidycreel)

# Four-week season: 10 weekdays, 8 weekend days sampled
calendar <- data.frame(
  date = as.Date(c(
    # Weekdays
    "2024-07-01", "2024-07-02", "2024-07-03", "2024-07-04", "2024-07-05",
    "2024-07-09", "2024-07-10", "2024-07-11", "2024-07-12", "2024-07-16",
    # Weekends
    "2024-07-06", "2024-07-07", "2024-07-13", "2024-07-14",
    "2024-07-20", "2024-07-21", "2024-07-27", "2024-07-28"
  )),
  day_type = c(rep("weekday", 10), rep("weekend", 8))
)

design <- creel_design(calendar, date = date, strata = day_type)
design
```

### Scheduling Circuit Start Times

Before field work begins, randomise the circuit start time for each survey day
using `generate_progressive_start()`. This ensures the count is unbiased with
respect to time-of-day effort patterns. Two strategies are available:

- `"discrete"` (default): start drawn from the \eqn{K = T / \tau} valid
  τ-aligned offsets — avoids mid-day over-representation caused by the common
  \eqn{U[0, T - \tau]} error. Requires \eqn{K} to be a whole number, and aborts
  if it is not; adjust `circuit_time` or the open hours to make it one.
- `"wraparound"`: start drawn from \eqn{U[0, T)}; circuit may wrap past the
  end of the survey day. Carries **no** divisibility requirement.

The whole-number requirement belongs to `"discrete"` alone, because that scheme
partitions the day into \eqn{K} non-overlapping blocks and samples one. Under
`"wraparound"` the count window slides freely, so \eqn{K = T/\tau} is an ordinary
ratio rather than a count of blocks and need not be a whole number.

This is why `add_counts()` does not check it. Both schemes give every moment of the
day equal probability of being observed, so both are estimated by the same
\eqn{\hat{E}_d = C \times T_d}; a divisibility check in the estimator would reject
perfectly valid wraparound designs. The constraint is enforced where it applies —
at scheduling time, by `generate_progressive_start()`.

```{r prog-start, message = FALSE}
starts <- generate_progressive_start(
  open_start    = "06:00",
  open_end      = "16:00",
  circuit_time  = 2,          # τ = 2 h; T = 10 h → k = 5 valid starts
  strategy      = "discrete",
  n             = nrow(calendar),
  seed          = 42
)
starts
```

The returned `creel_schedule` records `circuit_start`, `circuit_end`,
`is_wrapped`, and `direction` (`"forward"` / `"reverse"`) for each survey day.
Record the scheduled start and direction in your field protocol — both are
required for unbiased estimation.

### Count Data

Each row is one circuit traversal per sampled day. The `shift_hours` column records
the actual open hours for that day (here always 10, but could vary if a site closed
early due to weather).

```{r count-data, message = FALSE}
set.seed(7)
counts <- data.frame(
  date = calendar$date,
  day_type = calendar$day_type,
  n_anglers = c(
    # Weekday counts: moderate activity
    18L, 22L, 15L, 12L, 25L, 20L, 17L, 14L, 23L, 19L,
    # Weekend counts: higher activity
    48L, 55L, 42L, 61L, 53L, 47L, 58L, 64L
  ),
  shift_hours = 10
)
```

### Attaching Progressive Counts

Specify `count_type = "progressive"`, the circuit time $\tau$ in hours, and the
column holding $T_d$:

```{r add-counts, message = FALSE}
design <- add_counts(
  design, counts,
  count_type = "progressive",
  circuit_time = 2,
  period_length_col = shift_hours
)

# Per-day expanded effort (C × T_d) stored in design$counts
head(design$counts, 4)
```

The `n_anglers` column now holds $\hat{E}_d = C \times T_d$, not the raw count.
For the first day (18 anglers × 10 hours = 180 angler-hours):

```{r verify-formula}
# Verify: C × T_d = 18 × 10 = 180
design$counts$n_anglers[1]
```

### Effort Estimation

```{r effort, message = FALSE}
effort <- estimate_effort(design)
effort$estimates
```

The `estimate` column is the stratified total angler-hours for the 4-week season.
The `se` column is the standard error of that total; `se_within` is always zero for
progressive surveys because there is only one circuit per day (no within-day replication).

### Adding Interviews

Interviews are collected during or after the circuit traversal. Attach them exactly
as in any other survey type:

```{r interviews, message = FALSE}
set.seed(7)
n_int <- 120 # interviews collected across the season

int_dates <- sample(calendar$date, n_int, replace = TRUE)
catch_total <- rpois(n_int, lambda = 2.1)
interviews <- data.frame(
  date = int_dates,
  day_type = calendar$day_type[match(int_dates, calendar$date)],
  trip_status = "complete",
  hours_fished = round(pmax(rnorm(n_int, mean = 3.8, sd = 1.3), 0.5), 1),
  catch_total = catch_total,
  catch_kept = pmin(rpois(n_int, lambda = 0.7), catch_total)
)

design <- add_interviews(
  design, interviews,
  trip_status = trip_status,
  catch = catch_total,
  effort = hours_fished,
  n_anglers = 1, # every interview is a single angler
  harvest = catch_kept
)
```

### Catch Rate and Total Catch

```{r catch, message = FALSE}
cpue <- estimate_catch_rate(design)
cpue$estimates

total_catch <- estimate_total_catch(design)
total_catch$estimates
```

`estimate_total_catch()` multiplies the effort estimate by the CPUE estimate and
propagates uncertainty via the delta method, producing a season-total catch with
a combined standard error.

### Harvest Rate and Total Harvest

```{r harvest, message = FALSE}
harvest_rate <- estimate_harvest_rate(design)
harvest_rate$estimates

total_harvest <- estimate_total_harvest(design)
total_harvest$estimates
```

## Pope et al. Worked Example

Pope et al. (in press) give the canonical calculation: $C = 234$ anglers encountered
during a $\tau = 2$ h circuit on a day with $T_d = 8$ open hours.

$$\hat{E}_d = 234 \times 8 = 1{,}872 \text{ angler-hours}$$

```{r pope, message = FALSE}
cal_pope <- data.frame(
  date     = as.Date(c("2024-06-01", "2024-06-02")),
  day_type = c("weekday", "weekday")
)
des_pope <- creel_design(cal_pope, date = date, strata = day_type)

cnt_pope <- data.frame(
  date        = as.Date(c("2024-06-01", "2024-06-02")),
  day_type    = c("weekday", "weekday"),
  n_anglers   = c(234L, 200L),
  shift_hours = c(8, 8)
)
des_pope <- add_counts(
  des_pope, cnt_pope,
  count_type = "progressive",
  circuit_time = 2,
  period_length_col = shift_hours
)

# First-day Ê_d = 234 × 8 = 1,872 angler-hours
des_pope$counts
```

## Multiple Periods per Day

Some progressive surveys run two circuits per day (e.g., morning and evening). This
creates multiple counts per PSU (day), and tidycreel decomposes variance into between-day
and within-day components via the Rasmussen (1994) two-stage formula.

For the morning–evening case, use `count_time_col` to identify the circuit within each day:

```{r multi-period, message = FALSE}
# Two circuits per day (morning and evening traversals)
cal_2p <- data.frame(
  date     = rep(as.Date(c("2024-07-01", "2024-07-02", "2024-07-06", "2024-07-07")), 1),
  day_type = c("weekday", "weekday", "weekend", "weekend")
)
des_2p <- creel_design(cal_2p, date = date, strata = day_type)

cnt_2p <- data.frame(
  date = rep(as.Date(c(
    "2024-07-01", "2024-07-02",
    "2024-07-06", "2024-07-07"
  )), each = 2),
  day_type = rep(c("weekday", "weekday", "weekend", "weekend"), each = 2),
  count_time = rep(c("am", "pm"), 4),
  n_anglers = c(22L, 18L, 20L, 24L, 55L, 48L, 62L, 58L)
)

# Note: count_time_col for within-day identification
des_2p <- add_counts(des_2p, cnt_2p, count_time_col = count_time)

est_2p <- estimate_effort(des_2p)
est_2p$estimates
```

The `se_within` column is now non-zero, reflecting variability between the morning and
evening circuits within each day.

This chunk shows the within-day variance decomposition on its own; it does not apply the
progressive $T_d$ expansion, so `est_2p` is in counts rather than angler-hours. To get
both, add `count_type = "progressive"`, `circuit_time`, and `period_length_col` as in the
sections above — `add_counts()` supports multiple circuits per day on the progressive
path.

## Assemble a Summary Report

```{r summary, message = FALSE}
summary_tbl <- season_summary(list(
  effort  = effort,
  catch   = total_catch,
  harvest = total_harvest
))

summary_tbl$table
```

## References

- Hoenig, J. M., Robson, D. S., Jones, C. M., and Pollock, K. H. (1993). Scheduling
  counts in the instantaneous and progressive count methods for estimating sportfishing
  effort. *North American Journal of Fisheries Management*, 13, 723–736.

- Pope, K. L., Wilde, G. R., and Gabelhouse, D. W. Jr. (in press). Creel Surveys.
  Chapter 17 in *Fisheries Techniques*, 4th ed. American Fisheries Society, Bethesda, MD.

- Rasmussen, P. W. (1994). Two-stage variance estimation for creel surveys.

- Su, Y.-S., and Liu, P. (2025). Flexible creel survey estimators.
  *Canadian Journal of Fisheries and Aquatic Sciences*, 82, 1–27.
