| Title: | Parse Compact Numbers in Chinese Data |
| Version: | 0.2.5 |
| Description: | Parses compact numeric values, qualified quantities, and ranges commonly found in Chinese tables and spreadsheets. It handles Chinese magnitude suffixes, currencies, percentages, full-width characters, financial negatives, and configurable missing-value markers while reporting values that cannot be parsed safely. |
| License: | MIT + file LICENSE |
| URL: | https://github.com/Jorungandr/cncleanr |
| BugReports: | https://github.com/Jorungandr/cncleanr/issues |
| Encoding: | UTF-8 |
| Language: | en-US |
| Suggests: | spelling, testthat (≥ 3.0.0) |
| Config/testthat/edition: | 3 |
| Config/roxygen2/version: | 8.1.0 |
| NeedsCompilation: | no |
| Packaged: | 2026-09-21 04:37:14 UTC; ASUS |
| Author: | Haotian Liu [aut, cre, cph] |
| Maintainer: | Haotian Liu <wyvern06@163.com> |
| Repository: | CRAN |
| Date/Publication: | 2026-09-30 09:30:02 UTC |
cncleanr: Parse Compact Numbers in Chinese Data
Description
cncleanr parses compact numeric values commonly found in Chinese tables
and spreadsheets. It favors explicit rules and visible parse failures over
guessing.
Author(s)
Maintainer: Haotian Liu wyvern06@163.com [copyright holder]
Authors:
Haotian Liu wyvern06@163.com [copyright holder]
See Also
Useful links:
Extract Problems From a cncleanr Result
Description
Returns parsing failures attached to an object produced by a cncleanr parser. Successful results return an empty data frame with the same stable schema.
Usage
cn_problems(x)
Arguments
x |
An object returned by a cncleanr parsing function. |
Value
A data frame with columns index, value, and reason.
Examples
result <- suppressWarnings(parse_cn_number(c("bad", "2\u4e07")))
cn_problems(result)
Parse Compact Numbers Used in Chinese Data
Description
Converts character values such as "1.25\u4e07", "\uffe53\u4ebf",
"1.2e5", and "12.5%"
to doubles. Parsing is deliberately conservative: every non-missing value
must match the supported syntax in full. Spaces between digits are accepted
only when they form valid three-digit grouping.
Usage
parse_cn_number(
x,
na = c(
"", "NA", "N/A", "\u6682\u65e0", "\u672a\u516c\u5e03",
"\u2014", "\u2013", "-", "...", "\u2026"
),
strict = FALSE
)
Arguments
x |
A character, factor, or numeric vector. Numeric input is converted
to double and otherwise passed through unchanged, including |
na |
A character vector containing values that should be interpreted as missing. Matching happens after whitespace and full-width normalization. |
strict |
A single logical value. If |
Details
Finite-range validation applies to parsed character and factor
input. Existing numeric Inf, -Inf, and NaN values are preserved.
Value
A double vector with the same length and names as x. When invalid
values occur in non-strict mode, the result has a problems attribute with
columns index, value, and reason.
Examples
parse_cn_number(c("1.25\u4e07", "3\u4ebf\u5143", "12.5%", "\u6682\u65e0"))
parse_cn_number(c("\uff11\uff12\uff0e\uff15\uff05", "(2.5\u4e07)"))
Parse Qualified Quantities Used in Chinese Data
Description
Parses exact values and values qualified by language or symbols, such as
"\u7ea63\u4e07", "\u5927\u4e8e2\u4ebf", "\u4e0d\u5c0f\u4e8e5\u4e07", "10\u4e07+", "50\u4f59", or
"50\u4f59\u4e07\u5143".
The qualifier is retained instead of silently discarded.
Usage
parse_cn_quantity(
x,
na = c(
"", "NA", "N/A", "\u6682\u65e0", "\u672a\u516c\u5e03",
"\u2014", "\u2013", "-", "...", "\u2026"
),
strict = FALSE
)
Arguments
x |
A character, factor, or numeric vector. Numeric input is converted
to double and otherwise passed through unchanged, including |
na |
A character vector containing values that should be interpreted as missing. Matching happens after whitespace and full-width normalization. |
strict |
A single logical value. If |
Value
A data frame with double column value and character column
qualifier. Qualifiers are exact, approx, greater_than, at_least,
less_than, and at_most. Numeric finite and infinite values have the
qualifier exact; NA and NaN have a missing qualifier. Numeric values,
including Inf, -Inf, and NaN, are preserved in value.
Examples
parse_cn_quantity(c(
"\u7ea63\u4e07", "\u5927\u4e8e2\u4ebf", "10\u4e07+", "50\u4f59", "50\u4f59\u4e07\u5143"
))
Parse Numeric Ranges Used in Chinese Data
Description
Converts closed ranges such as "3\u4e07-5\u4e07" and open bounds such as
"10\u4e07\u5143\u4ee5\u4e0a" into explicit lower and upper bounds. A shared suffix in
"3-5\u4e07" is applied to both endpoints. Signs used by negative numbers and
scientific notation are distinguished from range separators.
Usage
parse_cn_range(
x,
na = c(
"", "NA", "N/A", "\u6682\u65e0", "\u672a\u516c\u5e03",
"\u2014", "\u2013", "-", "...", "\u2026"
),
strict = FALSE
)
Arguments
x |
A character, factor, or numeric vector. Numeric input is converted
to double and otherwise passed through unchanged, including |
na |
A character vector containing values that should be interpreted as missing. Matching happens after whitespace and full-width normalization. |
strict |
A single logical value. If |
Value
A data frame with columns lower, upper, lower_inclusive, and
upper_inclusive. Numeric input becomes an inclusive point range,
including numeric Inf and -Inf; NA and NaN have missing
inclusivity. Infinite bounds created from textual inequalities represent
open-ended bounds and are distinct from numeric point inputs.
Examples
parse_cn_range(c("3\u4e07-5\u4e07", "3-5\u4e07", "10\u4e07\u5143\u4ee5\u4e0a"))