Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

url

import url

This module provides classes and functions for parsing, building, resolving and processing URLs (and URL-like identifiers, such as mailto: and tel: links) per RFC 3986.

The scope of this module is not limited to HTTP or any single scheme: it will happily parse ftp://, ssh://, mailto:, or any scheme it has never heard of, and only makes scheme-specific decisions (like a default port) where the scheme is one this module actually recognizes.

Constructing a URL

%> import url
%> var link = url.Url('https', 'example.com', 9000)
%> link.absolute_url()
'https://example.com:9000'

Parsing a URL

The parse() function turns a URL string into a Url instance.

%> var link = url.parse('https://example.com:9000/path?q=1#frag')
%> link.scheme
'https'
%> link.host
'example.com'
%> link.port
9000
%> link.path
'/path'
%> link.query
'q=1'
%> link.hash
'frag'

Resolving relative URLs

resolve() implements RFC 3986’s relative-reference resolution algorithm, the same logic a browser uses to turn a relative href on a page into an absolute link.

%> var base = url.parse('https://example.com/a/b/c')
%> base.resolve('../d').to_string()
'https://example.com/a/d'

Query strings

query_params() decodes a URL’s query string into a dictionary of name -> [values] (a name can legally appear more than once in a query string, so every value is always a list).

%> url.parse('https://example.com?a=1&a=2&b=3').query_params()
{a: [1, 2], b: [3]}

The url API

Every public name in url, wherever it is declared. Each links to the page that documents it.

NameKindSummary
url.UrlclassThe Url class represents a parsed (or hand-built) URL and provides methods for inspecting, normalizing,…
url.UrlMalformedErrorclassRaised by parse() when strict is true and the given string cannot be interpreted as a well-formed URL.
url.decodefunctionDecodes URL-encoded string.
url.encodefunctionURL-encodes a string.
url.has_authorityfunctionReturns true if scheme (case-insensitively) allows the :// authority form directly; false for opaque…
url.parsefunctionParses given url string into a Url object.
url.parse_queryfunctionDecodes a raw query string (without a leading ?) into a dictionary mapping each parameter name to a list of…
url.remove_dot_segmentsfunctionRemoves . and .. path segments from path, per [RFC 3986…

Functions

remove_dot_segments()

url.remove_dot_segments(path: string) -> string

Removes . and .. path segments from path, per RFC 3986 §5.2.4. This is what keeps resolve() from leaking a base URL’s discarded segments into its result, e.g. /a/b/../c normalizes to /a/c.

Parameters

  • path (string)

Returns string

encode()

url.encode(url: string, strict: ?bool) -> string

URL-encodes a string.

This function is convenient when encoding a string to be used in a query part of a URL, as a convenient way to pass variables to the next page. Every byte of the string’s UTF-8 encoding that isn’t one of the reserved/unreserved characters defined by RFC 3986 is percent-encoded, so this is safe to use on text containing multi-byte characters.

If strict mode is enabled, space character is encoded with the percent (%) sign in order to conform with RFC 3986. Otherwise, is is encoded with the plus (+) sign in order to align with the default encoding used by modern browsers.

Parameters

  • url (string)
  • strict (?bool) — Default value is false

Returns string

decode()

url.decode(url: string) -> string

Decodes URL-encoded string. This function decodes any %## encoding in the given string and plus symbols (‘+’) to a space character.

Parameters

  • url (string)

Returns string

parse_query()

url.parse_query(query: string) -> dict

Decodes a raw query string (without a leading ?) into a dictionary mapping each parameter name to a list of its values. Equivalent to parse('?' + query).query_params(), but usable directly on a query string with no surrounding url.

%> url.parse_query('a=1&a=2&b=3')
{a: [1, 2], b: [3]}

Parameters

  • query (string)

Returns dict

has_authority()

url.has_authority(scheme: string) -> bool

Returns true if scheme (case-insensitively) allows the :// authority form directly; false for opaque schemes such as mailto or tel that address their content by an opaque string instead of a host.

Parameters

  • scheme (string)

Returns bool

parse()

url.parse(url: string, strict: ?bool) -> Url

Parses given url string into a Url object. If the strict argument is set to true, the parser will raise an UrlMalformedError when it encounters a malformed url; otherwise it makes a best effort and leaves whatever it couldn’t confidently parse as nil.

Parameters

  • url (string)
  • strict (?bool) — Default value is false

Returns Url

Raises UrlMalformedError if strict is true and the url is malformed.

Classes

UrlMalformedError

class url.UrlMalformedError < Error

Raised by parse() when strict is true and the given string cannot be interpreted as a well-formed URL.

Constructor

url.UrlMalformedError(message)

Parameters

  • message (string)

Url

class url.Url

The Url class represents a parsed (or hand-built) URL and provides methods for inspecting, normalizing, resolving and re-serializing it.

  • printable — has a @to_string(), so echo and print() show something useful
  • serializable — has a @to_json(), so it can be handed straight to json.encode()

Fields

FieldTypeDescription
schemeThe url scheme e.g. http, https, ftp, tcp etc. nil for a schemeless reference such as a bare path.
hostThe host information contained in the url.
portThe port contained in the url, as a number.
pathThe path of the URL.
hashThe url’s fragment/hash, without the leading #.
queryThe url’s raw query string, without the leading ?.
usernameThe username portion of the url’s userinfo, if any.
passwordThe password portion of the url’s userinfo, if any.
has_slashtrue if the url contains the :// (or bare //) authority marker.
empty_pathtrue if path is / only because no path was actually present in the original url (it was implied);…

Constructor

url.Url(scheme: ?string, host: ?string, port: ?number, path: ?string, query: ?string, hash: ?string, username: ?string, password: ?string, has_slash: ?bool, empty_path: ?bool)

Parameters

  • scheme (?string)
  • host (?string)
  • port (?number)
  • path (?string)
  • query (?string)
  • hash (?string)
  • username (?string)
  • password (?string)
  • has_slash (?bool)
  • empty_path (?bool)

Url.clone()

url.Url.clone() -> Url

Returns a clone of the current url.

Returns Url

Url.is_absolute()

url.Url.is_absolute() -> bool

Returns true if the url is absolute (has a scheme), false if it is a relative reference (e.g. a bare path).

Returns bool

Url.default_port()

url.Url.default_port() -> ?number

Returns the standard, well-known port for the url’s scheme (e.g. 80 for http, 443 for https), or nil if the scheme is unset or not one this module recognizes.

Returns ?number

Url.effective_port()

url.Url.effective_port() -> ?number

Returns port if the url specifies one, otherwise falls back to default_port().

Returns ?number

Url.query_params()

url.Url.query_params() -> dict

Decodes the url’s query string into a dictionary mapping each parameter name to a list of its values, since a query string can legally repeat the same name more than once. A name with no = (e.g. ?flag) decodes to a single empty-string value.

%> url.parse('https://example.com?a=1&a=2&b=3').query_params()
{a: [1, 2], b: [3]}

Returns dict

Url.get_param()

url.Url.get_param(name: string, fallback) -> any

Returns the first value of query parameter name, or fallback if it isn’t present. A convenience shortcut over query_params() for the common case of a parameter that only appears once.

Parameters

  • name (string)
  • fallback (any) — Default value is nil.

Returns any

Url.normalize()

url.Url.normalize() -> Url

Returns a new url with . and .. path segments resolved away, per RFC 3986 §5.2.4. Every other field is left untouched.

%> url.parse('https://example.com/a/b/../c').normalize().path
'/a/c'

Returns Url

Url.resolve()

url.Url.resolve(reference) -> Url

Resolves reference (a relative or absolute url string, or another Url instance) against this url as a base, per RFC 3986 §5, the same algorithm a browser uses to turn a page’s relative hrefs into absolute links.

%> var base = url.parse('https://example.com/a/b/c')
%> base.resolve('../d').to_string()
'https://example.com/a/d'
%> base.resolve('/x/y').to_string()
'https://example.com/x/y'
%> base.resolve('?q=1').to_string()
'https://example.com/a/b/c?q=1'

Parameters

  • reference (string|Url)

Returns Url

Url.authority()

url.Url.authority() -> string

Returns the url authority.

The authority component is preceded by a double slash (“//”) and is terminated by the next slash (“/”), question mark (“?”), or number sign (“#”) character, or by the end of the URI.

Returns string

Note: mailto and other opaque schemes have no authority. For this reason, they return an empty string as authority.

Url.host_for_display()

url.Url.host_for_display() -> string

Returns host, wrapped in [ ] brackets when it looks like an IPv6 literal, matching how it must appear when embedded back into a url string. Plain hostnames and IPv4 addresses are returned as-is.

Returns string

Url.host_is_ipv4()

url.Url.host_is_ipv4() -> bool

Returns true if the host of the url is a valid ipv4 address and false otherwise.

Returns bool

Url.host_is_ipv6()

url.Url.host_is_ipv6() -> bool

Returns true if the host of the url is a valid ipv6 address and false otherwise.

Returns bool

Url.absolute_url()

url.Url.absolute_url() -> string

Returns absolute url string of the url object.

Returns string

Url.to_string()

url.Url.to_string() -> string

Returns a string representation of the url object. This will only be the same as the absolute url if the original string is an absolute url.

Returns string

Url.equals()

url.Url.equals(other) -> bool

Returns true if other is a Url whose string form is identical to this one’s. Note that == does not call this: it always falls back to identity comparison for instances, so use a.equals(b) rather than a == b to compare two urls by value.

Parameters

  • other (any)

Returns bool


2021, Richard Ore and Zuri contributors