tidycreel.connect: Companion Package for Database Integration

What is tidycreel.connect?

tidycreel.connect is a companion package that bridges tidycreel with relational and embedded database back-ends. It targets agencies and research groups that collect creel survey data through field applications or custom data entry systems and need a reproducible, validated path to get that data into a form that tidycreel can analyse.

The package focuses on three layers of the data pipeline that tidycreel itself deliberately leaves out of scope:

  1. Connection management — opening validated connections to DBI-backed databases (SQL Server, DuckDB, SQLite), CSV file collections, or REST API endpoints with a consistent interface.
  2. Schema-aware data loading — reading counts, interviews, catch, and length records from the connected source, renaming raw fields to canonical tidycreel column names, and coercing types at import time.
  3. Survey discovery — listing and searching available creel surveys on API-backed connections before pulling data.

Installation

tidycreel.connect lives in a subdirectory of the tidycreel repository rather than a repository of its own, so both installers need to be told where to look:

# Using remotes
remotes::install_github("chrischizinski/tidycreel", subdir = "tidycreel.connect")

# Or using pak
pak::pak("chrischizinski/tidycreel/tidycreel.connect")

It is not on CRAN, and it requires tidycreel 7.0.0 or newer: identifier columns are normalised to character on both sides of the boundary, so an older tidycreel would hand the design numeric ids for the same survey.

How tidycreel.connect relates to tidycreel

tidycreel handles design, estimation, and reporting: building a creel_design object, running effort and catch estimators, producing publication-ready plots and tables.

tidycreel.connect handles the data-ingestion and storage layer that feeds it: connecting to a field database or API, loading validated data frames, and handing them off to tidycreel estimation functions.

The two packages are designed to work together but can be used independently:

A typical integrated workflow looks like this:

library(tidycreel)
library(tidycreel.connect)

# 1. Connect to a DBI-backed database (SQL Server, DuckDB, SQLite)
con <- creel_connect(dbi_connection, schema = "creel")

# -- or connect using a YAML configuration file --
con <- creel_connect_from_yaml("config/creel.yml")

# 2. Discover available surveys (API connections only)
list_creels(con)
search_creels(con, keyword = "2025")

# 3. Load validated data and hand off to tidycreel
counts     <- fetch_counts(con)
interviews <- fetch_interviews(con)

design <- creel_design(counts = counts, ...) |>
  add_interviews(interviews)

estimate_effort(design)

Exported functions

Connection management

Data loading

These functions use S3 dispatch so they work identically across DBI, CSV, and API back-ends. Each renames raw fields to canonical tidycreel column names and coerces data types at import time.

Survey discovery (API connections only)

CSV back-end

tidycreel.connect also supports a pure-CSV workflow for users who store survey data as flat files rather than in a database. Pass a named list of file paths to creel_connect():

con <- creel_connect(
  list(
    counts     = "data/counts.csv",
    interviews = "data/interviews.csv",
    catch      = "data/catch.csv"
  )
)

counts     <- fetch_counts(con)
interviews <- fetch_interviews(con)

More information