Skip to contents

This function returns the cleaned version of the URLs after applying protocol, www, case, and trailing slash handling rules. By default the result is a normalized canonical key composed of scheme, host, and path only; port is dropped (port_handling = "exclude"), and fragment/userinfo are always excluded (use get_port, get_fragment, or get_userinfo for those).

Usage

get_clean_url(
  url,
  protocol_handling = "keep",
  www_handling = "none",
  source = c("all", "private", "icann"),
  case_handling = "lower_host",
  trailing_slash_handling = "none",
  index_page_handling = "keep",
  path_normalization = "none",
  scheme_relative_handling = "keep",
  subdomain_levels_to_keep = NULL,
  host_encoding = "keep",
  path_encoding = "keep",
  query_handling = c("drop", "filter", "allow", "keep"),
  params_keep = NULL,
  params_drop = NULL,
  params_case_sensitive = FALSE,
  sort_params = FALSE,
  empty_param_handling = c("keep", "drop"),
  decode_plus = FALSE,
  port_handling = c("exclude", "keep", "strip_default", "strip_all"),
  scheme_policy = c("infer", "require"),
  scheme_acceptance = c("web", "general"),
  url_standard = NULL,
  engine = NULL,
  profile = NULL
)

Arguments

url

A character vector containing URLs to be parsed.

protocol_handling

A character string specifying how to handle protocols. Defaults to "keep". Regardless of this option, rurl only processes authority-based URLs whose scheme is one of http, https, ftp, or ftps; a scheme-bearing input with any other scheme (e.g. mailto:, tel:, ws:) yields parse_status = "error". Scheme inference (below) also requires the input to be host-shaped: a scheme-less string that is not a host (e.g. "asdfghjkl", "12345", "/path") or is a non-canonical IP literal (integer/hex/octal/short forms, or leading-zero octets like "192.168.010.1") is rejected as "error" rather than having a scheme fabricated for it.

  • "keep": If a supported scheme exists (http, https, ftp, ftps), it's used. If no scheme and the input is host-shaped, "http://" is added; otherwise the input is not a URL and yields "error".

  • "none": If a supported scheme exists, it's used. If no scheme, then no scheme is used (scheme component will be NA).

  • "strip": Any existing scheme is removed (scheme component will be NA).

  • "http": The scheme is forced to be "http".

  • "https": The scheme is forced to be "https".

www_handling

A character string specifying how to handle "www" and www[number] prefixes in the host. Defaults to "none".

  • "none": (Default) Leaves the host's www prefix (or lack thereof) untouched.

  • "strip": Removes any "www." or www[number]. prefix.

  • "keep": Ensures the host starts with "www.". If it has www[number]., it's normalized to "www.". If no www prefix, "www." is added. An empty input host remains empty.

  • "if_no_subdomain": If the host is a bare registered domain (e.g., "example.com"), "www." is added. If the host already has a "www." or www[number]. prefix, it is normalized to "www." (e.g., "www1.example.com" becomes "www.example.com"; "www1.sub.example.com" becomes "www.sub.example.com"). If a non-www subdomain exists (e.g., "sub.example.com" or the normalized "www.sub.example.com"), the host is not further altered. An empty input host remains empty.

source

Which PSL source to use: "all", "private", or "icann". Subdomain trimming depends on which section is consulted, so pass source = "icann" to exclude private suffixes (e.g. github.io).

case_handling

A character string specifying how to handle the case of the cleaned URL. Defaults to "lower_host", the RFC 3986 §6.2.2.1 normalization (scheme and host are case-insensitive and folded to lowercase; the path is case-sensitive and preserved).

  • "lower_host": (Default) Lowercases scheme and host only; the path keeps its original casing.

  • "keep": Preserves casing of the reconstructed URL.

  • "lower": Converts the cleaned URL to lowercase.

  • "upper": Converts the cleaned URL to uppercase.

trailing_slash_handling

A character string specifying how to handle trailing slashes in the path component of the cleaned URL. Defaults to "none".

  • "none": (Default) No specific handling is applied. Path remains as is after initial parsing.

  • "keep": Ensures a trailing slash. If a path exists and doesn't end with one, it's added. If path is just "/", it's kept.

  • "strip": Removes a trailing slash if present, unless the path is solely "/".

index_page_handling

A character string specifying how to handle index/default pages. Defaults to "keep".

  • "keep": (Default) Leave index/default page segments untouched.

  • "strip": Remove a trailing index.* or default.* segment (case-insensitive).

path_normalization

How to normalize path structure. Defaults to "none". rurl owns dot-segment resolution: the path is read from the input verbatim (not from libcurl's pre-normalized path), so "none" preserves . / .. segments (/a/../b stays /a/../b) and only the settings below change them. Resolution follows RFC 3986 section 5.2.4 and acts on literal ./.. segments only — a percent-encoded %2e is a normal path byte, never a dot segment, so it is never treated as traversal.

  • "none": (Default) No normalization; dot and slash structure is preserved exactly as written.

  • "collapse_slashes": Collapse duplicate slashes in the path.

  • "dot_segments": Resolve . and .. segments per RFC 3986.

  • "both": Apply both collapse_slashes and dot_segments.

scheme_relative_handling

How to handle URLs starting with "//". Defaults to "keep".

  • "keep": Parse using http but return scheme as NA and set status to "ok-scheme-relative".

  • "http": Assume http for parsing and output.

  • "https": Assume https for parsing and output.

  • "error": Treat scheme-relative URLs as invalid.

subdomain_levels_to_keep

An integer or NULL. Determines how many levels of subdomains are kept, in addition to any 'www.' prefix handled by www_handling.

  • NULL: (Default) No specific subdomain stripping is performed beyond www_handling.

  • 0: All subdomains are stripped. If www_handling preserved or added 'www.', it remains (e.g., 'www.sub.example.com' becomes 'www.example.com'; 'sub.example.com' becomes 'example.com').

  • N > 0: Keeps up to N levels of subdomains, counted from right-to-left (closest to the registered domain), in addition to any 'www.' prefix. E.g., if N=1, 'three.two.one.example.com' becomes 'one.example.com'; 'www.three.two.one.example.com' (post www_handling) becomes 'www.one.example.com'.

host_encoding

How to present the host in clean_url. Defaults to "keep".

  • "keep": Leave host as parsed by curl (may preserve original case).

  • "idna": Convert Unicode host labels to Punycode (IDNA) for the cleaned URL.

  • "unicode": Decode Punycode labels to Unicode for the cleaned URL.

path_encoding

How to present the path percent-encoding in clean_url — the readable-vs-browser rendering choice (the path analog of host_encoding). Defaults to "keep". This is an orthogonal presentation knob: it is independent of url_standard and layers on top of any profile (e.g. url_standard = "whatwg", path_encoding = "encode" emits the WHATWG-parsed path in browser form), exactly like host_encoding. Only "keep" preserves a profile's canonical identity path verbatim; "encode" and "decode" are presentation forms that may re-encode or decode reserved octets (so %2F may fold to a path-separating /), independent of whether a profile is set.

  • "keep": Leave the path percent-encoding untouched (the path is preserved as written in the URL, so %2F stays %2F rather than decoding into a path-separating /). With no url_standard, rurl keeps its historical RFC-style percent-hex case canonicalization, so %2f becomes %2F. Under url_standard = "whatwg", existing percent-triplet spelling is preserved byte-for-byte. Use "encode" to additionally normalize which bytes are encoded.

  • "encode": The browser/percent-encoded rendering. Decodes the path first, then percent-encodes each segment (slashes preserved), so a readable non-ASCII path is emitted in its percent-encoded UTF-8 form.

  • "decode": The readable rendering. Percent-decodes UTF-8 sequences in the path, so a percent-encoded segment is shown as readable text.

query_handling

A character string controlling whether (and how) the query string is included in clean_url. Defaults to "drop", which preserves the historical query-free clean_url. The raw query result field is never affected by this option — it always reports the faithful original query.

  • "drop": (Default) clean_url carries no query, exactly as before.

  • "filter": Keep contentful params, dropping known trackers via a built-in denylist (e.g. utm_*, fbclid, gclid). params_drop extends the denylist; params_keep rescues names (winning over both the denylist and empty-dropping).

  • "allow": Keep only params whose names match params_keep; all others are dropped. Here params_keep is an inclusion criterion only, not an empty-rescue.

  • "keep": Keep every param, re-encoded into canonical form (not the verbatim original — that stays on the query field).

In every non-"drop" mode the surviving query is re-encoded canonically (uppercase percent-hex, spaces as %20) and appended after the path. The query is intentionally EXEMPT from case_handling (query values are case-sensitive — tokens, IDs, signatures), so under case_handling = "lower" or "upper" the clean_url is no longer uniformly cased: scheme/host/path fold but the query keeps its original case. Because clean_url is the canonical_join key, any non-"drop" mode also brings the query into that join key (so ?id=1 and ?id=2 stop collapsing, while utm-only differences still collapse under "filter").

params_keep

Character vector of parameter-name globs (only * is special), or NULL (default). In "filter" mode this is the rescue list; in "allow" mode it is the allowlist. Ignored in "drop"/"keep".

params_drop

Character vector of parameter-name globs to add to the built-in denylist in "filter" mode, or NULL (default). Ignored in "drop"/"allow"/"keep".

params_case_sensitive

Logical (default FALSE). Controls whether the denylist and params_keep/params_drop matching is case-sensitive.

sort_params

Logical (default FALSE). When TRUE, surviving params are stably sorted by decoded key. Active in "filter"/"allow"/"keep".

empty_param_handling

One of "keep" (default) or "drop". "drop" removes empty-valued params (e.g. ?ref=), except those rescued by params_keep in "filter" mode.

decode_plus

Logical (default FALSE). When TRUE, + in query values is treated as a space (HTML-form decoding) before percent-decoding. FALSE keeps + literal (RFC 3986 generic behavior).

port_handling

A character string controlling whether the port appears in clean_url. Defaults to "exclude", today's only historical behavior. This knob is standalone and standard-independent (editorial, like www_handling) – url_standard never governs whether it may be set.

  • "exclude": (Default) The port never appears in clean_url.

  • "strip_all": Explicit alias of "exclude".

  • "keep": Include the syntactic port when present, including a default port under url_standard = "whatwg". This is an explicit non-parity override for callers that need the input's port spelling.

  • "strip_default": Keep only non-default ports (using the same scheme-default table), independent of url_standard.

scheme_policy

Controls whether scheme-less, host-shaped input is accepted (an input-acceptance axis, distinct from protocol_handling, which only controls how the scheme is presented, and from url_standard, which controls interpretation). Defaults to "infer".

  • "infer": (Default) Fabricate http:// for scheme-less host-shaped input (e.g. example.com parses as http://example.com), a browser-omnibox-style affordance. This is the historical behavior.

  • "require": Reject scheme-less input — a scheme-less host-shaped value becomes parse_status = "error" rather than gaining a fabricated scheme. Use this for a strict, pure-parser posture. Note this governs only bare host input; scheme-relative //host input is governed separately by scheme_relative_handling.

scheme_acceptance

Which scheme tokens may enter parsing (a scheme-acceptance axis, distinct from scheme_policy, which governs scheme-less input, and from url_standard, which governs interpretation). Defaults to "web".

  • "web": (Default) Only the curated web-scheme allowlist (http/https/ftp/ftps/file) is admitted; a scheme-bearing input outside it is parse_status = "error". This is the historical, byte-for-byte compatible behavior.

  • "general": Admit any syntactically valid scheme token and parse opaque (mailto:x), non-special (foo://host), and RFC-generic URLs. Requires an explicit url_standard ("rfc3986" or "whatwg"), which decides the interpretation; general with url_standard = NULL is an error. Non-special / opaque hosts receive no www-stripping, no domain/TLD derivation, and are never run through the IDNA/punycode helpers. A non-special scheme with no // is an opaque path: it has no authority, so host, user, port and the domain/tld columns are all NA and the entire remainder is the path (query/fragment are still split off). This includes mailto: — the recipient's @ never re-triggers authority parsing. To decompose a mailto: recipient, use the accessors (get_host() / get_domain() / get_user(), ADR 0012 D7) or get_mailto_recipients(); those deliberately return a recipient's parts where this table presents NA, because a recipient domain is extraction metadata, not the URL's authority.

url_standard

Optional top-level standard profile: NULL (default), "rfc3986", or "whatwg". With NULL the behavior is exactly what the individual low-level options select (fully backward compatible). When set, it selects a coherent set of standard-conformant behaviors for the axes it governs — path percent/dot handling, the host IPv4/reg-name model, and case_handling — so callers do not have to hand-assemble the low-level knobs. Passing a governed low-level knob (path_normalization or case_handling) with a value the selected profile would not choose is an error; passing the value the profile would pick is accepted (only case_handling = "lower_host" is accepted under a selector — "keep", "lower", and "upper" all conflict, since "lower" also lowercases the path, which neither standard sanctions). Added as the last argument so existing positional calls keep their meaning; always pass it by name. Under "whatwg" the selector additionally recognizes a literal backslash as a path separator for WHATWG-special schemes (http/https/ftp) and nulls default ports in parse output; use port_handling = "strip_default" for spec-style clean URL port rendering. See resolve_url for url_standard-governed reference resolution. The selector does not govern whether port_handling may be set (it is a standalone editorial knob), nor does it govern path_encoding (an orthogonal path-presentation knob that layers on any profile), IDNA rendering, or query handling.

engine

Optional pslr engine controlling which Public Suffix List backs domain / TLD / subdomain extraction: NULL (default) resolves against pslr's session-global default list — exactly the historical behavior — while a pslr::psl_engine() snapshot resolves against that specific list, per request, without mutating any global state (never call pslr::psl_use() for this). Use it to pin a particular list version or to load an alternate list via pslr::psl_engine(source = "path", path = ...). Process-local: an engine holds a C++ external pointer that does not serialize across R sessions or parallel workers — build it in the process that uses it; never cache it to disk or send it to a worker (rebuild one per process instead). Only the domain-derived outputs (domain, tld, and the subdomain-trimmed host / clean_url) depend on it.

profile

Optional named profile bundling several knobs at once: NULL (default; behaves exactly as the individual arguments select, fully backward compatible), "browser", "whatwg", "rfc-syntax", "seo", or the "seo" alias "canonical". A profile is separate from url_standard (it bundles acceptance, interpretation, leniency, and canonicalization together) and expands only into arguments you did not supply explicitly — an explicit argument always overrides the profile. "browser" is a browser-like fix-up posture (http-prepending; not Chrome-faithful); "whatwg" is the absolute-URL no-base posture that rejects scheme-less input (unlike a bare url_standard = "whatwg"); "rfc-syntax" is RFC 3986 generic syntax as parsing, not normalization (case and dot-segments are preserved); "seo"/"canonical" is rurl's origin-cleaning intent (https, strip www / trailing slash / index page, filter tracking params). Inspect the resolved bundle with url_profile. Also accepted by canonical_join (forwarded through its ...).

Value

A character vector of cleaned URLs.

Details

The query string is dropped by default (query_handling = "drop"), so the historical scheme/host/path output is byte-identical. Pass query_handling = "keep", "filter", or "allow" (with the companion params_* / sort_params / empty_param_handling / decode_plus arguments) to retain a shaped query on the cleaned URL; the engine is the same one safe_parse_url and get_query use, so get_clean_url(u, query_handling = "filter") equals safe_parse_url(u, query_handling = "filter")$clean_url.

The port is included only when port_handling != "exclude"; see safe_parse_url for the full port_handling semantics.

Examples

get_clean_url("Example.COM/Path") # Default lower_host: host folds, path kept
#> [1] "http://example.com/Path"
get_clean_url(
  "Example.COM/Path",
  case_handling = "keep",
  trailing_slash_handling = "keep"
)
#> [1] "http://Example.COM/Path/"
get_clean_url(
  "Example.COM/Path/",
  case_handling = "upper",
  trailing_slash_handling = "strip"
)
#> [1] "HTTP://EXAMPLE.COM/PATH"
get_clean_url("http://example.com", www_handling = "strip")
#> [1] "http://example.com/"
get_clean_url(
  "http://deep.sub.domain.example.com/path",
  subdomain_levels_to_keep = 0
)
#> [1] "http://example.com/path"
# -> "http://example.com/path"
get_clean_url(
  "http://www.deep.sub.domain.example.com/path",
  subdomain_levels_to_keep = 1,
  www_handling = "strip"
)
#> [1] "http://domain.example.com/path"
# -> "http://domain.example.com/path"
get_clean_url(
  "http://www.deep.sub.domain.example.com/path",
  subdomain_levels_to_keep = 1,
  www_handling = "keep"
)
#> [1] "http://www.domain.example.com/path"
# -> "http://www.domain.example.com/path"
# Query dropped by default (byte-identical to earlier releases):
get_clean_url("http://example.com/p?utm_source=nl&id=42")
#> [1] "http://example.com/p"
# -> "http://example.com/p"
# Strip trackers, keep contentful params:
get_clean_url(
  "http://example.com/p?utm_source=nl&id=42",
  query_handling = "filter"
)
#> [1] "http://example.com/p?id=42"
# -> "http://example.com/p?id=42"