url
import url
This module provides classes and functions for parsing, building,
resolving and processing URLs (and URL-like identifiers, such as
mailto: and tel: links) per RFC
3986.
The scope of this module is not limited to HTTP or any single scheme: it
will happily parse ftp://, ssh://, mailto:, or any scheme it has
never heard of, and only makes scheme-specific decisions (like a default
port) where the scheme is one this module actually recognizes.
Constructing a URL
%> import url
%> var link = url.Url('https', 'example.com', 9000)
%> link.absolute_url()
'https://example.com:9000'
Parsing a URL
The parse() function turns a URL string into a Url instance.
%> var link = url.parse('https://example.com:9000/path?q=1#frag')
%> link.scheme
'https'
%> link.host
'example.com'
%> link.port
9000
%> link.path
'/path'
%> link.query
'q=1'
%> link.hash
'frag'
Resolving relative URLs
resolve() implements RFC 3986’s relative-reference resolution
algorithm, the same logic a browser uses to turn a relative href on a
page into an absolute link.
%> var base = url.parse('https://example.com/a/b/c')
%> base.resolve('../d').to_string()
'https://example.com/a/d'
Query strings
query_params() decodes a URL’s query string into a dictionary of name -> [values]
(a name can legally appear more than once in a query string, so every
value is always a list).
%> url.parse('https://example.com?a=1&a=2&b=3').query_params()
{a: [1, 2], b: [3]}
The url API
Every public name in url, wherever it is declared. Each links to the
page that documents it.
| Name | Kind | Summary |
|---|---|---|
url.Url | class | The Url class represents a parsed (or hand-built) URL and provides methods for inspecting, normalizing,… |
url.UrlMalformedError | class | Raised by parse() when strict is true and the given string cannot be interpreted as a well-formed URL. |
url.decode | function | Decodes URL-encoded string. |
url.encode | function | URL-encodes a string. |
url.has_authority | function | Returns true if scheme (case-insensitively) allows the :// authority form directly; false for opaque… |
url.parse | function | Parses given url string into a Url object. |
url.parse_query | function | Decodes a raw query string (without a leading ?) into a dictionary mapping each parameter name to a list of… |
url.remove_dot_segments | function | Removes . and .. path segments from path, per [RFC 3986… |
Functions
remove_dot_segments()
url.remove_dot_segments(path: string) -> string
Removes . and .. path segments from path, per RFC 3986
§5.2.4. This is what
keeps resolve() from leaking a base URL’s discarded segments into its
result, e.g. /a/b/../c normalizes to /a/c.
Parameters
path(string)
Returns string
encode()
url.encode(url: string, strict: ?bool) -> string
URL-encodes a string.
This function is convenient when encoding a string to be used in a query part of a URL, as a convenient way to pass variables to the next page. Every byte of the string’s UTF-8 encoding that isn’t one of the reserved/unreserved characters defined by RFC 3986 is percent-encoded, so this is safe to use on text containing multi-byte characters.
If strict mode is enabled, space character is encoded with the percent (%) sign in order to conform with RFC 3986. Otherwise, is is encoded with the plus (+) sign in order to align with the default encoding used by modern browsers.
Parameters
url(string)strict(?bool) — Default value isfalse
Returns string
decode()
url.decode(url: string) -> string
Decodes URL-encoded string. This function decodes any %## encoding in the given string and plus symbols (‘+’) to a space character.
Parameters
url(string)
Returns string
parse_query()
url.parse_query(query: string) -> dict
Decodes a raw query string (without a leading ?) into a dictionary
mapping each parameter name to a list of its values. Equivalent to
parse('?' + query).query_params(), but usable directly on a query
string with no surrounding url.
%> url.parse_query('a=1&a=2&b=3')
{a: [1, 2], b: [3]}
Parameters
query(string)
Returns dict
has_authority()
url.has_authority(scheme: string) -> bool
Returns true if scheme (case-insensitively) allows the ://
authority form directly; false for opaque schemes such as mailto or
tel that address their content by an opaque string instead of a host.
Parameters
scheme(string)
Returns bool
parse()
url.parse(url: string, strict: ?bool) -> Url
Parses given url string into a Url object. If the strict argument is set
to true, the parser will raise an UrlMalformedError when it
encounters a malformed url; otherwise it makes a best effort and leaves
whatever it couldn’t confidently parse as nil.
Parameters
url(string)strict(?bool) — Default value isfalse
Returns Url
Raises UrlMalformedError if strict is true and the url is
malformed.
Classes
UrlMalformedError
class url.UrlMalformedError < Error
Raised by parse() when strict is true and the given string cannot
be interpreted as a well-formed URL.
Constructor
url.UrlMalformedError(message)
Parameters
message(string)
Url
class url.Url
The Url class represents a parsed (or hand-built) URL and provides methods for inspecting, normalizing, resolving and re-serializing it.
- printable — has a
@to_string(), soechoandprint()show something useful - serializable — has a
@to_json(), so it can be handed straight tojson.encode()
Fields
| Field | Type | Description |
|---|---|---|
scheme | The url scheme e.g. http, https, ftp, tcp etc. nil for a schemeless reference such as a bare path. | |
host | The host information contained in the url. | |
port | The port contained in the url, as a number. | |
path | The path of the URL. | |
hash | The url’s fragment/hash, without the leading #. | |
query | The url’s raw query string, without the leading ?. | |
username | The username portion of the url’s userinfo, if any. | |
password | The password portion of the url’s userinfo, if any. | |
has_slash | true if the url contains the :// (or bare //) authority marker. | |
empty_path | true if path is / only because no path was actually present in the original url (it was implied);… |
Constructor
url.Url(scheme: ?string, host: ?string, port: ?number, path: ?string, query: ?string, hash: ?string, username: ?string, password: ?string, has_slash: ?bool, empty_path: ?bool)
Parameters
scheme(?string)host(?string)port(?number)path(?string)query(?string)hash(?string)username(?string)password(?string)has_slash(?bool)empty_path(?bool)
Url.clone()
url.Url.clone() -> Url
Returns a clone of the current url.
Returns Url
Url.is_absolute()
url.Url.is_absolute() -> bool
Returns true if the url is absolute (has a scheme), false if it is a
relative reference (e.g. a bare path).
Returns bool
Url.default_port()
url.Url.default_port() -> ?number
Returns the standard, well-known port for the url’s scheme (e.g. 80
for http, 443 for https), or nil if the scheme is unset or not
one this module recognizes.
Returns ?number
Url.effective_port()
url.Url.effective_port() -> ?number
Returns port if the url specifies one, otherwise falls back to
default_port().
Returns ?number
Url.query_params()
url.Url.query_params() -> dict
Decodes the url’s query string into a dictionary mapping each parameter
name to a list of its values, since a query string can legally repeat
the same name more than once. A name with no = (e.g. ?flag) decodes
to a single empty-string value.
%> url.parse('https://example.com?a=1&a=2&b=3').query_params()
{a: [1, 2], b: [3]}
Returns dict
Url.get_param()
url.Url.get_param(name: string, fallback) -> any
Returns the first value of query parameter name, or fallback if it
isn’t present. A convenience shortcut over query_params() for the
common case of a parameter that only appears once.
Parameters
name(string)fallback(any) — Default value isnil.
Returns any
Url.normalize()
url.Url.normalize() -> Url
Returns a new url with . and .. path segments resolved away, per RFC
3986 §5.2.4. Every other field is left untouched.
%> url.parse('https://example.com/a/b/../c').normalize().path
'/a/c'
Returns Url
Url.resolve()
url.Url.resolve(reference) -> Url
Resolves reference (a relative or absolute url string, or another
Url instance) against this url as a base, per RFC 3986
§5, the same algorithm a
browser uses to turn a page’s relative hrefs into absolute links.
%> var base = url.parse('https://example.com/a/b/c')
%> base.resolve('../d').to_string()
'https://example.com/a/d'
%> base.resolve('/x/y').to_string()
'https://example.com/x/y'
%> base.resolve('?q=1').to_string()
'https://example.com/a/b/c?q=1'
Parameters
reference(string|Url)
Returns Url
Url.authority()
url.Url.authority() -> string
Returns the url authority.
The authority component is preceded by a double slash (“//”) and is terminated by the next slash (“/”), question mark (“?”), or number sign (“#”) character, or by the end of the URI.
Returns string
Note: mailto and other opaque schemes have no authority. For this reason, they return an empty string as authority.
Url.host_for_display()
url.Url.host_for_display() -> string
Returns host, wrapped in [ ] brackets when it looks like an IPv6
literal, matching how it must appear when embedded back into a url
string. Plain hostnames and IPv4 addresses are returned as-is.
Returns string
Url.host_is_ipv4()
url.Url.host_is_ipv4() -> bool
Returns true if the host of the url is a valid ipv4 address and false otherwise.
Returns bool
Url.host_is_ipv6()
url.Url.host_is_ipv6() -> bool
Returns true if the host of the url is a valid ipv6 address and false otherwise.
Returns bool
Url.absolute_url()
url.Url.absolute_url() -> string
Returns absolute url string of the url object.
Returns string
Url.to_string()
url.Url.to_string() -> string
Returns a string representation of the url object. This will only be the same as the absolute url if the original string is an absolute url.
Returns string
Url.equals()
url.Url.equals(other) -> bool
Returns true if other is a Url whose string form is identical to
this one’s. Note that == does not call this: it always falls back to
identity comparison for instances, so use a.equals(b) rather than a == b
to compare two urls by value.
Parameters
other(any)
Returns bool
2021, Richard Ore and Zuri contributors