Creel surveys are often paired with tagging studies to estimate either angler population size or seasonal exploitation rate. tidycreel provides three estimators for these purposes:
estimate_angler_n() —
closed-population mark-recapture for total angler count (Chapman,
Petersen, or Schnabel)estimate_mr_harvest() — total harvest
derived from a mark-recapture population estimateestimate_exploitation_rate() —
seasonal exploitation rate from tagged-fish recoveries (Pollock et
al. 1994)All three return creel_estimates objects compatible with
the standard tidycreel output tools (print(),
autoplot(), write_estimates()).
| Goal | Estimator |
|---|---|
| Estimate total anglers from a tag-and-resight study | estimate_angler_n() |
| Convert angler population estimate to total harvest | estimate_mr_harvest() |
| Estimate fraction of population harvested from tagged-fish recoveries | estimate_exploitation_rate() |
The Chapman estimator is a bias-corrected Petersen estimator recommended when the recapture count is small. With \(M\) tagged anglers released, \(n\) anglers checked in the second sample, and \(m\) recaptures:
\[\hat{N} = \frac{(M+1)(n+1)}{(m+1)} - 1\]
Variance:
\[\widehat{\text{Var}}(\hat{N}) = \frac{(M+1)(n+1)(M-m)(n-m)}{(m+1)^2(m+2)}\]
library(tidycreel)
result_chapman <- estimate_angler_n(M = 200L, n = 50L, m = 10L)
print(result_chapman)
#>
#> ── Creel Survey Estimates ──────────────────────────────────────────────────────
#> Method: mark-recapture-chapman
#> Variance: chapman
#> Confidence level: 95%
#>
#> # A tibble: 1 × 6
#> parameter estimate se ci_lower ci_upper n
#> <chr> <dbl> <dbl> <dbl> <dbl> <int>
#> 1 N_hat 931. 232. 605. 1715. 10
The n column in the estimates tibble records the
recapture count (m), which determines precision.
The unadjusted Lincoln–Petersen estimator:
\[\hat{N} = \frac{M \cdot n}{m}\]
tidycreel enforces a minimum of \(m \geq 7\) recaptures; below this threshold the Petersen estimator carries large positive bias and Chapman should be used instead.
result_petersen <- estimate_angler_n(M = 200L, n = 50L, m = 10L, method = "petersen")
print(result_petersen)
#>
#> ── Creel Survey Estimates ──────────────────────────────────────────────────────
#> Method: mark-recapture-petersen
#> Variance: petersen
#> Confidence level: 95%
#>
#> # A tibble: 1 × 6
#> parameter estimate se ci_lower ci_upper n
#> <chr> <dbl> <dbl> <dbl> <dbl> <int>
#> 1 N_hat 1000 283. 605. 1715. 10
For \(K \geq 2\) successive sampling occasions, the Schnabel estimator pools information across occasions, carrying Chapman’s (1952) small-sample correction to the recapture total:
\[\hat{N} = \frac{\sum_{k=1}^{K} M_k n_k}{\sum_{k=1}^{K} m_k + 1}\]
where \(M_k\) is the cumulative count of marked-at-large anglers before occasion \(k\) (so \(M_1 = 0\)), \(n_k\) is the catch on occasion \(k\), and \(m_k\) is the number of recaptures.
The \(+1\) is the default as of
tidycreel 3.0.0. Each \(m_k\) is
approximately Poisson with parameter \(M_k n_k
/ N\), which is what motivates the correction. Dettloff (2023)
simulated both forms and found the unadjusted estimator turns biased
high at moderate sample sizes — inflating harvest downstream —
while the adjusted form’s bias approaches zero without ever becoming
positive, at lower variance and no cost in large samples. It also keeps
bias handling consistent across occasion counts: Schnabel reduces to
Lincoln-Petersen at \(K = 2\), so an
unadjusted Schnabel would have silently dropped the correction that
method = "chapman" applies to the same data. Pass
bias_adjust = FALSE for the pre-3.0.0 form, which is also
what fishmethods::schnabel() computes.
The correction is worth \(-1/(\sum m_k + 1)\) in relative terms: \(-33\%\) at \(\sum m_k = 2\), \(-1.9\%\) at 52, and \(-0.2\%\) at 500. It matters at low recapture totals and is negligible at high ones.
Confidence intervals use the Poisson distribution when \(\sum m_k < 50\) and a normal
approximation on \(1/\hat{N}\)
otherwise. The Poisson branch inverts the distribution of \(\sum m_k\) directly rather than centring on
\(\hat{N}\), so it is a valid interval
under either form and its bounds do not move with
bias_adjust; the normal branch is built around \(1/\hat{N}\) and does.
result_schnabel <- estimate_angler_n(
M = c(0L, 47L, 91L, 131L),
n = c(50L, 50L, 50L, 50L),
m = c(0L, 4L, 6L, 8L),
method = "schnabel"
)
print(result_schnabel)
#>
#> ── Creel Survey Estimates ──────────────────────────────────────────────────────
#> Method: mark-recapture-schnabel
#> Variance: delta
#> Confidence level: 95%
#>
#> # A tibble: 1 × 6
#> parameter estimate se ci_lower ci_upper n
#> <chr> <dbl> <dbl> <dbl> <dbl> <int>
#> 1 N_hat 708. 158. 498. 1345 18
Four sampling occasions yield 18 total recaptures here, so the Poisson CI branch is used automatically.
The same data can be pooled by regression instead of by ratio. Schumacher and Eschmeyer (1943) fit \(m_k/n_k\) against \(M_k\) through the origin with slope \(1/N\), which Seber (1982, sec. 4.1.3) develops as
\[\hat{N} = \frac{\sum_{k} n_k M_k^2}{\sum_{k} m_k M_k}\]
with the interval from his eq. (4.17) on \(K - 2\) degrees of freedom. The degrees of freedom differ from Schnabel’s \(K - 1\) for a stated reason: \(y_1\) is identically zero when \(M_1 = 0\), so it is not a random observation and does not enter the regression. It needs at least three occasions.
result_se <- estimate_angler_n(
M = c(0L, 47L, 91L, 131L),
n = c(50L, 50L, 50L, 50L),
m = c(0L, 4L, 6L, 8L),
method = "schumacher"
)
print(result_se)
#>
#> ── Creel Survey Estimates ──────────────────────────────────────────────────────
#> Method: mark-recapture-schumacher
#> Variance: delta
#> Confidence level: 95%
#>
#> # A tibble: 1 × 6
#> parameter estimate se ci_lower ci_upper n
#> <chr> <dbl> <dbl> <dbl> <dbl> <int>
#> 1 N_hat 699. 44.7 548. 964. 18
Which of the two to use is a design question. Seber expects the regression form “to be robust with regard to departures from the underlying assumptions” and recommends it “in conjunction with the other methods” — as a cross-check rather than a replacement. Dettloff (2023) found the two adjusted forms effectively equivalent at larger sample sizes, with Schumacher-Eschmeyer less variable and Schnabel reaching unbiasedness slightly sooner.
Do not choose whichever gives the narrower interval. Hansen and Van Kirk (2018) computed both and selected the estimator producing the smallest 95% CI, and that procedure does not have 95% coverage — picking the narrower of two intervals after seeing them conditions on the luckier draw. tidycreel does not implement the rule. Decide before looking, or report both.
Once angler population size is estimated,
estimate_mr_harvest() scales it by a known harvest rate to
obtain total harvest. The harvest rate is typically derived from creel
interview data — the mean number of fish kept per angler. It is a rate
in fish per angler, not a proportion, and values above 1 are
ordinary.
\[\hat{H} = \hat{N} \times r, \quad SE(\hat{H}) = r \times SE(\hat{N})\]
# harvest_rate derived from creel interviews: 0.35 fish kept per angler
harvest <- estimate_mr_harvest(angler_n = result_chapman, harvest_rate = 0.35)
#> ℹ `harvest_rate_se` was not supplied, so the reported SE excludes harvest-rate
#> uncertainty.
#> It reflects uncertainty in the abundance estimate alone and is a lower bound.
#> Supply `harvest_rate_se` to propagate the rate's own error (Goodman 1960).
print(harvest)
#>
#> ── Creel Survey Estimates ──────────────────────────────────────────────────────
#> Method: mark-recapture-harvest
#> Variance: delta
#> Confidence level: 95%
#>
#> # A tibble: 1 × 5
#> parameter estimate se ci_lower ci_upper
#> <chr> <dbl> <dbl> <dbl> <dbl>
#> 1 total_harvest 326. 81.1 212. 600.
estimate_angler_n() reports its unit as NA
rather than "anglers", and
estimate_mr_harvest() inherits that. The estimators never
see what the marking protocol marked — anglers, boats, or parties — so
neither can say which of those \(\hat{N}\) counts.
The delta method propagates uncertainty in \(\hat{N}\) only. Uncertainty in the harvest rate itself is not propagated in this release; if harvest rate uncertainty is substantial, re-run across plausible bounds as a sensitivity check.
estimate_exploitation_rate() implements the Pollock et
al. (1994) moment estimator. Rather than counting anglers, it estimates
the fraction of a tagged cohort that was harvested during the season.
The inputs come from two sources:
estimate_total_harvest()) with standard error \(SE_C\). This must be harvest, not catch —
released fish stayed in the tagged cohort, so a catch total from
estimate_total_catch() inflates \(\hat{u}\).\[\hat{u} = \frac{C \cdot m}{T \cdot n}, \quad \widehat{\text{Var}}(\hat{u}) \approx \left(\frac{C}{T}\right)^2 \frac{p(1-p)}{n} + p^2 \frac{SE_C^2}{T^2}\]
where \(p = m/n\).
result_expl <- estimate_exploitation_rate(
T = 200L,
C = 450.0,
se_C = 42.0,
n = 180L,
m = 15L
)
#> ℹ `C` was supplied as a bare number, so its target cannot be verified.
#> It must be a season (period) total -- a sampled-day total understates `u` by
#> the sampling fraction.
#> Pass the `estimate_total_harvest()` result itself to have this checked.
print(result_expl)
#>
#> ── Creel Survey Estimates ──────────────────────────────────────────────────────
#> Method: Exploitation Rate (Mark-Recapture)
#> Variance: delta
#> Confidence level: 95%
#> Unit: proportion
#>
#> # A tibble: 1 × 8
#> estimate se ci_lower ci_upper n T C m
#> <dbl> <dbl> <dbl> <dbl> <int> <int> <dbl> <int>
#> 1 0.188 0.0495 0.0897 0.285 180 200 450 15
When the survey spans multiple strata (day types, access areas),
supply a data frame with one row per stratum. The T-weighted aggregate
is automatically appended as an .overall row:
\[\hat{u}_{overall} = \frac{\sum_h T_h \hat{u}_h}{\sum_h T_h}, \quad \widehat{\text{Var}}(\hat{u}_{overall}) = \frac{\sum_h T_h^2 \,\widehat{\text{Var}}(\hat{u}_h)}{(\sum_h T_h)^2}\]
strata_df <- data.frame(
stratum = c("weekday", "weekend"),
T_h = c(120L, 80L),
C_h = c(280.0, 170.0),
se_C_h = c(28.0, 22.0),
n_h = c(110L, 70L),
m_h = c(9L, 6L)
)
result_strat <- estimate_exploitation_rate(strata = strata_df, by = "stratum")
print(result_strat)
#>
#> ── Creel Survey Estimates ──────────────────────────────────────────────────────
#> Method: Exploitation Rate (Mark-Recapture)
#> Variance: delta
#> Confidence level: 95%
#> Grouped by: stratum
#> Unit: proportion
#>
#> # A tibble: 3 × 9
#> stratum estimate se ci_lower ci_upper n T C m
#> <chr> <dbl> <dbl> <dbl> <dbl> <int> <int> <dbl> <int>
#> 1 weekday 0.191 0.0639 0.0643 0.318 110 120 280 9
#> 2 weekend 0.182 0.0749 0.0327 0.332 70 80 170 6
#> 3 .overall 0.187 0.0487 0.0914 0.283 180 200 450 15
Pass aggregate = FALSE to suppress the
.overall row.
If not all harvested tagged fish are reported, supply the reporting rate \(\lambda \in (0, 1]\):
\[\hat{u}_{adj} = \frac{\hat{u}}{\lambda}\]
result_adj <- estimate_exploitation_rate(
T = 200L,
C = 450.0,
se_C = 42.0,
n = 180L,
m = 15L,
reporting_rate = 0.80
)
#> ℹ `C` was supplied as a bare number, so its target cannot be verified.
#> It must be a season (period) total -- a sampled-day total understates `u` by
#> the sampling fraction.
#> Pass the `estimate_total_harvest()` result itself to have this checked.
print(result_adj)
#>
#> ── Creel Survey Estimates ──────────────────────────────────────────────────────
#> Method: Exploitation Rate (Mark-Recapture)
#> Variance: delta
#> Confidence level: 95%
#> Unit: proportion
#>
#> # A tibble: 1 × 8
#> estimate se ci_lower ci_upper n T C m
#> <dbl> <dbl> <dbl> <dbl> <int> <int> <dbl> <int>
#> 1 0.234 0.0619 0.112 0.357 180 200 450 15
Uncertainty in \(\lambda\) is not propagated; treat reporting rate accuracy as a separate sensitivity analysis.
| Scenario | Recommended estimator |
|---|---|
| Single recapture event, small \(m\) (\(< 7\)) | estimate_angler_n(method = "chapman") |
| Single recapture event, \(m \geq 7\) | "chapman" or "petersen" (Chapman
preferred) |
| Multiple survey occasions with cumulative marking | estimate_angler_n(method = "schnabel") |
| Three or more occasions, wanting a regression cross-check | estimate_angler_n(method = "schumacher") |
| Converting population estimate to total harvest | estimate_mr_harvest() |
| Fraction of tagged cohort harvested (season-level) | estimate_exploitation_rate() |
All three estimators share the standard closed-population assumptions:
For estimate_exploitation_rate() two additional caveats
apply:
Hansen, M. J., & Van Kirk, R. W. (2018). A mark-recapture-based approach for estimating angler harvest. North American Journal of Fisheries Management, 38(2), 400–410. https://doi.org/10.1002/nafm.10038
Pollock, K. H., Jones, C. M., & Brown, T. L. (1994). Angler Survey Methods and Their Applications in Fisheries Management. AFS Special Publication 25. American Fisheries Society.
Jones, C. M., & Pollock, K. H. (2012). Recreational survey methods: estimating effort, harvest, and abundance. In A. V. Zale et al. (Eds.), Fisheries Techniques (3rd ed., Ch. 19). American Fisheries Society.
Chapman, D. G. (1952). Inverse, multiple and sequential sample censuses. Biometrics, 8(4), 286–306. https://doi.org/10.2307/3001864
Dettloff, K. (2023). Assessment of bias and precision among simple closed population mark-recapture estimators. Fisheries Research, 265, 106756. https://doi.org/10.1016/j.fishres.2023.106756
Schumacher, F. X., & Eschmeyer, R. W. (1943). The estimation of fish populations in lakes or ponds. Journal of the Tennessee Academy of Science, 18, 228–249.
Seber, G. A. F. (1982). The Estimation of Animal Abundance and Related Parameters, 2nd ed. Macmillan, New York.
De Lury, D. B. (1958). The estimation of population size by a marking and recapture procedure. Journal of the Fisheries Research Board of Canada, 15(1), 19–25. https://doi.org/10.1139/f58-003