The explicit set of `rurl` canonicalization arguments that determine a pagerankr node identity (scheme + host + path). Every knob that *shapes that key* is pinned here with an explicit value, so node keys never depend on `rurl`'s own defaults – which have changed across `rurl` versions (e.g. `case_handling` flipped from `"keep"` to `"lower_host"`) and previously desynced the pagerankr <-> semantic join. Knobs that only govern components the key drops (port, query, fragment, userinfo) are deliberately left unpinned – see the "Knobs deliberately left unpinned" note in @details.
Details
The canonical node key is **scheme + host + path**, with the path percent-decoded and RFC 3986 dot-segments removed (port, query, fragment and userinfo are dropped by `get_clean_url`). These ten arguments fix exactly how that key is derived.
The pinned values are chosen to reproduce that committed canonical key, not to match `rurl`'s defaults. As of `rurl` 2.1.0 they intentionally **override** two defaults: `path_normalization = "dot_segments"` (default `"none"`) and `path_encoding = "decode"` (default `"keep"`). `rurl` 2.1.0 redefined its `"none"`/`"keep"` path defaults to keep the path verbatim, which silently changed the key and desynced the pagerankr <-> semantic join for paths containing dot-segments or percent-encoding; pinning the explicit values above restores the original keys. The remaining knobs still equal `rurl`'s current defaults. (The anti-drift guarantee ultimately lives in semantic's byte-parity oracle test, which caught this; the pin is a convenience that must be re-pinned when a dependency redefines a value.)
The same ten arguments are accepted by both `rurl::get_clean_url()` (the cleaning path) and `rurl::safe_parse_url()` (the domain-filtering path), so one profile drives both and the two paths stay symmetrical.
**Knobs deliberately left unpinned.** `rurl::get_clean_url()` has grown options that govern components the node key does not carry, or that select a parsing route rather than a key component: * `source` (default `"all"`) – the public-suffix source, reachable only through `www_handling`/`subdomain_levels_to_keep`, both pinned to values that do not consult it. * `query_handling` (default `"drop"`) and its parameter sub-options (`params_keep`, `params_drop`, `params_case_sensitive`, `sort_params`, `empty_param_handling`, `decode_plus`) – the key drops the query. * `port_handling` (default `"exclude"`) – the key drops the port. * `url_standard` (default `NULL` – no standard profile, added in `rurl` 2.2.0) and, added in `rurl` 2.7.0, `scheme_policy` (default `"infer"`), `scheme_acceptance` (default `"web"`), `engine` (default `NULL`) and `profile` (default `NULL`, an unrelated `rurl` concept that merely shares a name with this function).
Under the scheme+host+path key these have no visible effect at their defaults, so pinning them would add noise without changing identity. They are intentionally **not** part of this profile; instead `test-canonicalization.R` guards them from two sides, so a `rurl` change is caught on the pagerankr side rather than silently changing node identity: a behavioral guard asserts that a canonical key really does drop the port, query and fragment (catching a default *flip*), and a surface guard reads `rurl`'s own formals and fails on any argument this profile has neither pinned nor listed above (catching an *addition*). A committed node-key probe covering the cross-platform parse-determinism risk surface pins the keys themselves.
The cross-repo contract requires **semantic** to pin the identical profile; change both repos together.
**Accepted divergence on un-canonicalizable input.** For a value `rurl` cannot parse (an unsupported scheme like `mailto:`/`tel:`, whitespace, a dotless bare token), `rurl` returns `NA`. pagerankr's [clean_url_columns()] keeps such a value as its raw self so it survives as an opaque graph node (see that function; PR #50), whereas semantic's `canonical_url()` returns `None` and drops it (FR-05 rurl byte-parity). This is intentional and does **not** break the `node_score` <-> `page` join: valid URLs still produce byte-identical keys on both sides (the actual contract), and in the semantic -> pagerankr bridge semantic canonicalizes and drops un-canonicalizable inputs *before* pagerankr sees the edges, so the raw fallback never fires on that path. It only affects pagerankr run standalone on raw crawl data, where such tokens become opaque nodes instead of being dropped.
Examples
# The pinned profile that determines pagerankr node identity.
profile <- canonical_profile()
str(profile)
#> List of 10
#> $ protocol_handling : chr "keep"
#> $ case_handling : chr "lower_host"
#> $ www_handling : chr "none"
#> $ trailing_slash_handling : chr "none"
#> $ index_page_handling : chr "keep"
#> $ path_normalization : chr "dot_segments"
#> $ scheme_relative_handling: chr "keep"
#> $ subdomain_levels_to_keep: NULL
#> $ host_encoding : chr "keep"
#> $ path_encoding : chr "decode"
# Key knobs that shape the scheme + host + path node key.
profile$case_handling # "lower_host"
#> [1] "lower_host"
profile$path_normalization # "dot_segments"
#> [1] "dot_segments"
profile$path_encoding # "decode"
#> [1] "decode"