---
title: "tidycreel and the survey Package: A Side-by-Side Guide"
output:
  rmarkdown::html_vignette:
    highlight: null
vignette: >
  %\VignetteIndexEntry{tidycreel and the survey Package: A Side-by-Side Guide}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include = FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>"
)
```

## Introduction

The tidycreel package is built on top of R's `survey` package. Every estimation
function in tidycreel — `estimate_effort()`, `estimate_catch_rate()`,
`estimate_total_catch()` — is a wrapper around a corresponding `survey` package
call. The design objects tidycreel constructs are standard `svydesign` objects
under the hood, and the estimators delegate directly to `svymean()`,
`svytotal()`, and `svyratio()`.

This vignette walks through the same effort + catch rate + total catch analysis
twice: first using raw `survey` package calls, then using tidycreel. The goal is
to make the mapping explicit, so that R users familiar with the `survey` package
can immediately understand what tidycreel is doing and why. If you already know
`svyratio()` and the delta method for variance propagation, you will recognize
those computations in tidycreel's output.

## Data Setup

Both workflows use the same three example datasets included with tidycreel. Load
them once and they are shared across both parts of this vignette.

```{r data-setup}
library(tidycreel)

data(example_calendar)
data(example_counts)
data(example_interviews)

head(example_calendar)
head(example_counts)
head(example_interviews)
```

The calendar covers 14 days (10 weekdays, 4 weekend days). The counts record
observed effort (angler-hours) on each sampled day. The interviews record
catch and effort for individual completed and incomplete trips.

---

## Part 1 — Raw survey Package Workflow

### 1a. Survey Design from a Stratified Calendar

A creel survey with weekday/weekend stratification is a stratified simple
random sample. The finite population correction (FPC) for each stratum is the
total number of days in that stratum over the survey period.

`example_counts` already carries a `day_type` column (copied from the
calendar), so we can build the FPC column directly from the calendar frequency
table.

```{r raw-design}
library(survey)

# Stratum sizes: total days per day_type in the calendar
stratum_sizes <- table(example_calendar$day_type)
stratum_sizes

# Attach FPC to the count frame
counts_frame <- example_counts
counts_frame$fpc <- as.integer(stratum_sizes[counts_frame$day_type])
counts_frame

# Build the stratified survey design
svy_counts <- svydesign(
  ids    = ~1,
  strata = ~day_type,
  fpc    = ~fpc,
  data   = counts_frame
)
svy_counts
```

Because all 14 survey days are observed (a census of the 14-day period), the
FPC reduces to 1 and the standard error of the totals is 0. In a real survey
where only a subset of days are sampled, the FPC would be less than 1 and
variances would be positive.

### 1b. Effort Estimation via svytotal

Total effort is the sum of observed angler-hours across the 14 days, estimated
via `svytotal()` on the count frame design.

```{r raw-effort}
effort_total <- svytotal(~effort_hours, svy_counts)
effort_total
confint(effort_total)
```

The result is the estimated total angler-hours for the survey period, with
a 95% confidence interval. Because this example uses a complete census of all
14 days, the standard error is 0. In practice, creel surveys sample a subset
of days, and `svytotal()` extrapolates from the observed days to the full
survey period using the stratum weights.

### 1c. Catch Rate via svyratio and Total Catch via the Delta Method

Catch per unit effort (CPUE) is estimated from the interview data using a
ratio-of-means estimator: the total catch across all complete interviews
divided by the total effort across those same interviews. This is the
`svyratio()` approach.

```{r raw-cpue}
# Use complete trips only (standard practice)
complete_trips <- subset(example_interviews, trip_status == "complete")
nrow(complete_trips)

# Build a simple survey design over the interview frame
svy_interviews <- svydesign(ids = ~1, data = complete_trips)

# Ratio estimator: total catch / total effort (ratio-of-means CPUE)
cpue_ratio <- svyratio(~catch_total, ~hours_fished, svy_interviews)
cpue_ratio
```

The CPUE estimate is the ratio of total catch to total effort across all
complete interviews. The `svyratio()` function computes the Taylor linearization
variance for this ratio, accounting for the correlation between catch and effort
within each interview.

Now combine the effort total and CPUE using the delta method to propagate
variance through the product E × C:

```{r raw-total-catch}
effort_coef <- coef(effort_total)[[1]]
cpue_coef <- coef(cpue_ratio)[[1]]
var_effort <- vcov(effort_total)[[1]]
var_cpue <- vcov(cpue_ratio)[[1]]

# Delta method: Var(effort * cpue) = effort^2 * Var(cpue) + cpue^2 * Var(effort)
total_catch_est <- effort_coef * cpue_coef
total_catch_var <- effort_coef^2 * var_cpue + cpue_coef^2 * var_effort
total_catch_se <- sqrt(total_catch_var)

cat("Total catch estimate:", round(total_catch_est, 1), "\n")
cat("Standard error:      ", round(total_catch_se, 1), "\n")
cat(
  "95% CI: [",
  round(total_catch_est - 1.96 * total_catch_se, 1), ",",
  round(total_catch_est + 1.96 * total_catch_se, 1), "]\n"
)
```

---

## Part 2 — tidycreel Equivalent

### 2a. creel_design() + add_counts()

`creel_design()` wraps the calendar and stratification into a design object.
`add_counts()` attaches the count frame and builds the internal `svydesign`
object — the equivalent of the `svydesign()` call in Part 1a above.

```{r tc-design}
design <- creel_design(example_calendar, date = date, strata = day_type)
design <- add_counts(design, example_counts)
design
```

The design reports the same 14 observations across two strata (weekday,
weekend) as the raw `svydesign` object above.

### 2b. estimate_effort()

`estimate_effort()` calls `svytotal()` on the internal design object and
returns the result as a tidy tibble with labelled columns.

```{r tc-effort}
effort_est <- estimate_effort(design)
effort_est
```

The estimate (372.5 angler-hours) matches the raw `svytotal()` result in
Part 1b. tidycreel adds the CI directly to the output and labels the column
as `estimate` instead of the raw column name.

### 2c. estimate_catch_rate() + add_interviews()

Before calling `estimate_catch_rate()`, attach the interview data via
`add_interviews()`. This maps interview columns to the design vocabulary
and filters to complete trips by default.

```{r tc-interviews}
design <- add_interviews(design, example_interviews,
  catch         = catch_total,
  effort        = hours_fished,
  n_anglers     = n_anglers,
  harvest       = catch_kept,
  trip_status   = trip_status,
  trip_duration = trip_duration
)
```

Now estimate CPUE, which calls `svyratio()` internally on the complete-trip
subset of the interview frame:

```{r tc-cpue}
cpue_est <- estimate_catch_rate(design)
cpue_est
```

The CPUE estimate (2.29) and SE (0.11) match the `svyratio()` output from
Part 1c.

### 2d. estimate_total_catch()

`estimate_total_catch()` applies the delta method to combine the effort and
CPUE estimates, exactly as in Part 1c:

```{r tc-total-catch}
total_catch <- estimate_total_catch(design)
total_catch
```

The total catch estimate (approximately 851) is consistent with the raw
delta-method calculation in Part 1c. Small numerical differences reflect
tidycreel's internal handling of the complete-trip filter and variance
components — the statistical method is identical.

---

## Mapping Table

| Step | survey package | tidycreel |
|---|---|---|
| Design construction | `svydesign(ids=~1, strata=~day_type, fpc=~fpc, data=counts)` | `creel_design()` + `add_counts()` |
| Attach interview data | subset + `svydesign(ids=~1, data=complete_trips)` | `add_interviews()` |
| Effort estimation | `svytotal(~effort_hours, design)` | `estimate_effort(design)` |
| Catch rate | `svyratio(~catch_total, ~hours_fished, int_design)` | `estimate_catch_rate(design)` |
| Total catch | E × C with delta method Var(E×C) = E² Var(C) + C² Var(E) | `estimate_total_catch(design)` |
| Variance method | Taylor linearization (default in survey) | Taylor linearization (default in tidycreel) |

---

## When to Use Each Approach

**Use tidycreel** for standard creel survey designs where the main analysis
follows the effort + CPUE + total catch workflow. tidycreel handles the
bookkeeping (FPC construction, complete-trip filtering, delta method variance)
and returns tidy tibbles ready for reporting. The domain vocabulary
(`creel_design`, `estimate_effort`, `estimate_catch_rate`) matches the
language creel biologists already use.

**Use raw survey package calls** when you need capabilities outside tidycreel's
scope: custom clustering structures, probability-proportional-to-size (PPS)
sampling, replicate-weight designs, or nonstandard variance estimators. The
`as_creel_svydesign()` function extracts the internal `svydesign` object from any
tidycreel design, giving you full access to the survey package toolbox while
still using tidycreel for the initial data setup.

```{r extract-design, eval = FALSE}
# Extract the internal svydesign object for advanced use
internal_svy <- as_creel_svydesign(design)
# Now use any survey package function directly
svymean(~effort_hours, internal_svy)
```
