String Methods
Every method on the built-in string type, with its signature, what it
returns, and the cases where it does something other than the obvious
thing.
| Method | Returns | Summary |
|---|---|---|
length() | number | Returns the length of a string. |
upper() | string | Returns a copy of the string with all the cased characters converted to uppercase. |
lower() | string | Return a copy of the string with all the cased characters converted to lowercase. |
is_alpha() | boolean | Returns true if all the characters in the string are all alphabets and the string is not empty., otherwise returns false. |
is_alnum() | boolean | Returns true if all the characters in the string are either alphabets or numbers and the string is not empty, otherwise returns false. |
is_number() | Returns true if all the characters in the string are all digits and the string is not empty, otherwise returns false. | |
is_lower() | boolean | Returns true if at least one character in the string is cased, all cased characters are lower cased and the string is not empty. |
is_upper() | boolean | Returns true if at least one character in the string is cased, all cased characters are upper cased and the string is not empty. |
is_space() | boolean | Returns true if there are only whitespace characters in the string and the string is not empty. |
ord() | number | Returns the Unicode code point of the string, which must be exactly one character long. |
trim(chars: ?string) | string | Returns a copy of the string with characters stripped from both ends. |
ltrim(chars: ?string) | string | Returns a copy of the string with characters stripped from its start only. |
rtrim(chars: ?string) | string | Returns a copy of the string with characters stripped from its end only. |
join(string: string) | string | Returns a string which is a concatenation of the items in the iterable using the string as the separator. |
split(delimiter: string) | list | Returns a list of words or characters in a string after separating the content of the string at every point where the delimiter is found. |
index_of(str: string, start_index: ?number) | number | Returns the index position of the first occurrence of the string str in the string string. |
last_index_of(str: string, end_index: ?number) | number | Returns the index position of the last occurrence of the string str in the string string, searching from the end. |
starts_with(str: string) | boolean | Returns true if the string begins with the string or character specified in str, otherwise it returns false. |
ends_with(str: string) | boolean | Returns true if the string ends with the string or character specified in str, otherwise it returns false. |
count(str: string) | number | Returns the number of non-overlapping occurrences of the substring str in the string. |
to_number(base) | number | Returns the first numeric value contained in the string if any exists or 0 if the string contains no numeric value. |
to_bigint(base) | bigint | Returns the integer value of the string as a bigint, or 0n if the string does not spell one. |
to_list() | list | Returns a list whose elements consists of every character contained in the string in order of appearance. |
to_bytes() | bytes | Returns the content of the string as a stream of bytes. |
lpad(width: number, fill: ?string) | string | Returns the string left justified in a string of length width. |
rpad(width: number, fill: ?string) | string | Returns the string right justified in a string of length width. |
match(str: string) | boolean|dictionary | If the string str is a regular string, this method returns true if the string contains a substring str. |
matches(reg: string) | dictionary | Returns a dictionary containing every match of the given regular expression reg in the source string. |
replace(str: string, replacement: string, use_regex: ?bool) | string | Returns a copy of the string with all occurrences or matches of str replaced by the replacement string. |
replace_with(regex: string, callback: function) | string | Returns a copy of the string with all occurrences or matches of regex replaced with the result of the function callback which is invoked only if and after a match has occurred. |
ascii() | string | Reinterprets the string as a raw byte view: each byte of its UTF-8 encoding becomes its own character (a codepoint between 0 and 255, i.e. |
case_fold() | string | Returns a copy of the string case-folded for case-insensitive comparison, using full Unicode case folding rather than plain lowercasing. |
compare(other: string) | number | Compares the string with another string. |
is_empty() | boolean | Returns true if the string is empty, false otherwise. |
contains(str: string) | boolean | Returns true if the string contains the specified substring, false otherwise. |
lines() | list | Returns the lines of the string as an list as it would be if split on newline characters. |
each_line(callback: function) | void | Iterates over each line of the string, calling the provided callback function with the line and its index. |
each(callback: function) | void | Iterates over each character of the string, calling the provided callback function with the character and its index. |
capitalize() | string | Returns a new string with the first character capitalized and the rest in lowercase. |
title() | string | Returns a new string with each word capitalized. |
to_string() | string | Returns the string itself. |
length()
length() -> number
Returns the length of a string. Note that this method is UTF-8
compatible and will return the UTF-8 length for the string if the string
contains UTF-8 characters whether written directly or via the \u or
\U escapes.
For example:
%> 'This is a pretty long string'.length()
28
%> 'उनका एक समय'.length()
11
%> 'This text mixes English and 粵語'.length()
30
Returns number
upper()
upper() -> string
Returns a copy of the string with all the cased characters converted to
uppercase. Note that the result of this method may return false when
tested with is_upper() of the string contains Unicode characters
that are not case folded.
For example:
%> 'zuri'.upper()
'ZURI'
Returns string
lower()
lower() -> string
Return a copy of the string with all the cased characters converted to
lowercase.
For example:
%> 'Zuri Is Bae'.lower()
'zuri is bae'
Returns string
is_alpha()
is_alpha() -> boolean
Returns true if all the characters in the string are all alphabets and
the string is not empty., otherwise returns false.
For example:
%> 'abracadabra'.is_alpha()
true
%> 'my tooth aches'.is_alpha()
false
%> ''.is_alpha()
false
Returns boolean
is_alnum()
is_alnum() -> boolean
Returns true if all the characters in the string are either alphabets
or numbers and the string is not empty, otherwise returns false. This
method is the same as string.is_alpha() or string.is_number().
For example:
%> '3Idiots'.is_alnum()
true
%> 'Three Idiots'.is_alnum()
false
%> '3 Idiots'.is_alnum()
false
%> '3'.is_alnum()
true
%> 'idiots'.is_alnum()
true
%> ''.is_alnum()
false
Returns boolean
is_number()
is_number()
Returns true if all the characters in the string are all digits and
the string is not empty, otherwise returns false.
For example:
%> '123.5'.is_number()
false
%> '1970'.is_number()
true
%> '1980s'.is_number()
false
is_lower()
is_lower() -> boolean
Returns true if at least one character in the string is cased, all
cased characters are lower cased and the string is not empty. Otherwise,
it returns false.
For example:
%> 'all'.is_lower()
true
%> 'all...123'.is_lower()
true
%> 'All...123'.is_lower()
false
%> ''.is_lower()
false
Returns boolean
is_upper()
is_upper() -> boolean
Returns true if at least one character in the string is cased, all
cased characters are upper cased and the string is not empty. Otherwise,
it returns false.
For example:
%> 'ALL'.is_upper()
true
%> 'ALL...123'.is_upper()
true
%> 'All...123'.is_upper()
false
%> ''.is_upper()
false
Returns boolean
is_space()
is_space() -> boolean
Returns true if there are only whitespace characters in the string and
the string is not empty. Otherwise, it returns empty.
For example:
%> '. '.is_space()
false
%> '\r\n'.is_space()
true
%> '\t '.is_space()
true
Returns boolean
ord()
ord() -> number
Returns the Unicode code point of the string, which must be exactly one character long.
%> 'A'.ord()
65
%> 'AB'.ord()
Unhandled Error: ord() must be called on a single character, got AB
StackTrace:
<repl>:1 -> @.script()
Returns number
Raises Error if the string is not exactly one character long.
trim()
trim(chars: ?string) -> string
Returns a copy of the string with characters stripped from both ends.
With no argument, whitespace is stripped: space, tab (\t), line feed
(\n), vertical tab, form feed and carriage return (\r). Other
Unicode spaces, such as a no-break space, are kept.
Given chars, every character in it is stripped instead, in any order
and any number of times, until a character not in chars is reached
at each end. chars is a set of characters, not a prefix or suffix:
'xyax'.trim('xy') is 'a'. An empty chars strips nothing.
The string itself is never changed. A string with nothing to strip comes back as an equal copy.
For example:
%> ' example '.trim()
'example'
%> '\t example \r\n'.trim()
'example'
%> ' example '.trim('e')
' example '
%> 'example'.trim('e')
'xampl'
%> '--==example==--'.trim('-=')
'example'
Parameters
chars(?string) — The characters to strip (Default = whitespace).
Returns string
ltrim()
ltrim(chars: ?string) -> string
Returns a copy of the string with characters stripped from its start
only. The characters stripped are chosen exactly as they are for
trim(): whitespace when chars is not given, or every character of
chars when it is.
For example:
%> ' example '.ltrim()
'example '
%> 'example'.ltrim('e')
'xample'
%> '0012'.ltrim('0')
'12'
Parameters
chars(?string) — The characters to strip (Default = whitespace).
Returns string
rtrim()
rtrim(chars: ?string) -> string
Returns a copy of the string with characters stripped from its end only.
The characters stripped are chosen exactly as they are for trim():
whitespace when chars is not given, or every character of chars
when it is.
For example:
%> ' example '.rtrim()
' example'
%> 'example'.rtrim('e')
'exampl'
%> 'line\r\n'.rtrim('\r\n')
'line'
Parameters
chars(?string) — The characters to strip (Default = whitespace).
Returns string
join()
join(string: string) -> string
Returns a string which is a concatenation of the items in the iterable using the string as the separator. If the iterable contains just one item or the string is empty, the original element is returned. If the iterable contains non-string items, the items are converted to their string representation before joining.
Bytes are the only non supported iterables.
For example:
%> ','.join(['ok', 1, true])
'ok,1,true'
%> '--'.join('name')
'n--a--m--e'
%> ','.join('a')
'a'
Parameters
string(string) — The string to join the items in the iterable.
Returns string
split()
split(delimiter: string) -> list
Returns a list of words or characters in a string after separating the content of the string at every point where the delimiter is found.
If the delimiter is an empty string, the resultant list will contain the individual characters of the string in the order in which they appear in the original string. Consecutive delimiters are not grouped together and are deemed to delimit empty strings. Splitting an empty string with a specified separator returns an empty list.
This method has full UTF-8 support.
For example:
%> 'name'.split('')
[n, a, m, e]
%> '1<>2<>3'.split('<>')
[1, , 2, , 3]
%> '1,2,3'.split(',')
[1, 2, 3]
%> ''.split(',')
[]
%> '地点'.split('')
[地, 点]
%> 'who is in the garden'.split('/\s/')
[who, is, in, the, garden]
Parameters
delimiter(string) — The delimiter to use the split the string.
Returns list
index_of()
index_of(str: string, start_index: ?number) -> number
Returns the index position of the first occurrence of the string str
in the string string. If the str cannot be found anywhere in
string, it returns -1. If the start_index parameter is given, it
will start scanning from the given index.
For example:
%> 'hello, world'.index_of(' ')
6
%> 'hello, world'.index_of('e')
1
%> 'hello, world'.index_of('q')
-1
%> 'hello, world'.index_of('o')
4
%> 'hello, world'.index_of('o', 5) # next index of `o` starting from index 5.
8
Parameters
str(string) — The string to search for.start_index(?number) — The index to start the search from.
Returns number
last_index_of()
last_index_of(str: string, end_index: ?number) -> number
Returns the index position of the last occurrence of the string str
in the string string, searching from the end. If str cannot be
found anywhere in string, it returns -1.
If the end_index parameter is given, only a match that begins at or
before that index counts. That is the same thing index_of()’s own
second parameter bounds, so for any index n, index_of(str, n) and
last_index_of(str, n) are the first and last matches of the two halves
n splits the string into.
An empty str returns -1, matching index_of().
For example:
%> 'hello, world'.last_index_of('o')
8
%> 'hello, world'.last_index_of('l')
10
%> 'hello, world'.last_index_of('q')
-1
%> 'hello, world'.last_index_of('o', 7) # last `o` starting at or before index 7.
4
Splitting a path on its final separator is the usual reason to reach for it:
%> var path = 'a/b/c'
%> path.last_index_of('/')
3
%> path[path.last_index_of('/') + 1, path.length()]
'c'
Parameters
str(string) — The string to search for.end_index(?number) — The highest index a match may start at.
Returns number
starts_with()
starts_with(str: string) -> boolean
Returns true if the string begins with the string or character
specified in str, otherwise it returns false.
For example:
%> 'hello, world'.starts_with('hello')
true
%> 'hello, world'.starts_with('hellios')
false
Parameters
str(string) — The string to search for.
Returns boolean
ends_with()
ends_with(str: string) -> boolean
Returns true if the string ends with the string or character specified
in str, otherwise it returns false.
For example:
%> 'gumtree'.ends_with('tree')
true
%> 'gumtree'.ends_with('mree')
false
Parameters
str(string) — The string to search for.
Returns boolean
count()
count(str: string) -> number
Returns the number of non-overlapping occurrences of the substring str in the string.
For those coming from Python who may consider this method similar to Python’s own, this method differs in that it does not allow specifying a start and end region for the operation. Zuri considers this unnecessary as the same can be accomplished by slicing the string.
For example:
%> 'Hallelujah'.count('l')
3
%> 'ding dong'.count('ng')
2
%> 'ding dong'[2,7].count('ng') # setting region to search for counts - 'ng do'
1
Parameters
str(string) — The string to search for.
Returns number
to_number()
to_number(base) -> number
Returns the first numeric value contained in the string if any exists or
0 if the string contains no numeric value. Floating numbers that have
the same value as their integer counterparts will return the integer
value.
For example:
%> '123.0 hell'.to_number()
123
%> '427 and 12'.to_number()
427
%> '96.3 of 31'.to_number()
96.3
%> 'error'.to_number()
0
Parameters
base(number) — The base the digits are in, from 2 to 36. Defaults to10. A fractional part is only read in base 10, since no other base spells one.
Returns number
Raises RangeError if base is outside 2 to 36.
to_bigint()
to_bigint(base) -> bigint
Returns the integer value of the string as a bigint, or 0n if the
string does not spell one.
This is to_number() for integers too large to be a number. A number is
exact only up to 2^53; past that, digits are lost, and an id or a
BIGINT UNSIGNED read from a database routinely runs past it. Every
digit survives here however long the run.
%> '9007199254740993'.to_bigint()
9007199254740993n
%> '9007199254740993'.to_number() # rounded down by one
9007199254740992
%> '-42'.to_bigint()
-42n
%> 'ff'.to_bigint(16)
255n
%> 'row 427 of 12'.to_bigint()
427n
%> 'error'.to_bigint()
0n
The number is found exactly as to_number() finds it: the first one
written in the string, with any text around it ignored. No fractional
part is read, since this produces an integer, so '12.5'.to_bigint() is
12n.
Parameters
base(number) — The base the digits are in, from 2 to 36. Defaults to10.
Returns bigint
Raises RangeError if base is outside 2 to 36.
to_list()
to_list() -> list
Returns a list whose elements consists of every character contained in
the string in order of appearance. Characters that repeat in the string
will have different entries in the same index as they appear in the
string.
For example:
%> 'Zuri'.to_list()
[Z, u, r, i]
%> 'Plantation'.to_list()
[P, l, a, n, t, a, t, i, o, n]
Returns list
to_bytes()
to_bytes() -> bytes
Returns the content of the string as a stream of bytes.
The Zuri REPL may truncate long bytes data when printing to console/terminal.
For example:
%> 'Zuri'.to_bytes()
(42 6c 61 64 65)
%> 'Plantation'.to_bytes()
(50 6c 61 6e 74 61 74 69 6f 6e)
Returns bytes
lpad()
lpad(width: number, fill: ?string) -> string
Returns the string left justified in a string of length width. Padding
is done using the specified character fill if given of a space (' ')
if a fill is not specified. The original string is returned if width
is less than string.length().
For example:
%> 'cat'.lpad(5)
' cat'
%> 'cat'.lpad(5, '-')
'--cat'
%> 'cat'.lpad(2, '-')
'cat'
Parameters
width(number) — The length of the string after padding.fill(?string) — The character to use for padding.
Returns string
rpad()
rpad(width: number, fill: ?string) -> string
Returns the string right justified in a string of length width.
Padding is done using the specified character fill if given of a space
(' ') if a fill is not specified. The original string is returned if
width is less than string.length().
For example:
%> 'Hmm'.rpad(6)
'Hmm '
%> 'Hmm'.rpad(6, '.')
'Hmm...'
%> 'Hmm'.rpad(3, '.')
'Hmm'
Parameters
width(number) — The length of the string after padding.fill(?string) — The character to use for padding.
Returns string
match()
match(str: string) -> boolean|dictionary
If the string str is a regular string, this method returns true if
the string contains a substring str. Otherwise, it returns false.
If the string str contains a valid regular
expression (we’ll get to that shortly below), it
returns false if a match for the regex str cannot be found in the
string. Otherwise, it returns a dictionary of the
first match: the whole match under key 0, each capture group under its
number, and a named group under its name as well. A group that takes no
part in the match is nil. Groups that share a name, under the J
modifier, give the name to the one that took part.
If the offset argument is specified, it becomes the offset in the string at which to start matching.
For example:
%> 'gorilla'.match('go') # regular string match
true
%> 'gorilla'.match('gox') # regular string non-match
false
%> 'gorilla'.match('/?gox/') # regular expression match
{0: go}
%> 'gorilla'.match('/gox\d/') # regular expression non-match
false
%> '2024-01'.match('/(?<year>\d+)-(\d+)/')
{0: 2024-01, 1: 2024, 2: 01, year: 2024}
Parameters
str(string) — The string to match.
Returns boolean|dictionary
matches()
matches(reg: string) -> dictionary
Returns a dictionary containing every match of the given regular expression reg in the source string. If no match is found, an empty dictionary is returned.
If the offset argument is specified, it becomes the offset in the string at which to start matching.
For example:
%> '123 dollars'.matches('/[a-z]+|\d+/')
{0: [123, dollars]}
%> 'who is in the garden'.matches('/\w+/')
{0: [who, is, in, the, garden]}
Parameters
reg(string) — The regular expression to match.
Returns dictionary
replace()
replace(str: string, replacement: string, use_regex: ?bool) -> string
Returns a copy of the string with all occurrences or matches of str replaced by the replacement string.
In the replacement string, if str is a regular expression, then
capture groups can be referenced using the syntax $index. Taking as an
example, capture group 0 contains the entire match and can be used in
the replacement string as $0.
To escape the
$sign in the replacement string, use the double backslashes (\\).
For example:
%> 'lady friend'.replace('d', 'z') # non-regex
'lazy frienz'
%> 'John is 26 years old'.replace('/(\d+)/', '1$1') # regex example
'John is 126 years old'
%> 'John is 26 years old'.replace('/(\d+)/', '1\\$2')
'John is 1$2 years old'
Parameters
str(string) — The string to match.replacement(string) — The replacement string.use_regex(?bool) — Whether to use the regular expression or the string string as the match string (default = true).
Returns string
Note: When the third parameter
use_regexis set to false, str will never be treated as a regular expression even if it contains a valid regular expression.
replace_with()
replace_with(regex: string, callback: function) -> string
Returns a copy of the string with all occurrences or matches of regex replaced with the result of the function callback which is invoked only if and after a match has occurred.
The callback function is defined as follows:
def replacer(match, p1, p2, /* …, */ pN, offset, string) {
return replacement
}
The arguments to the function are as follows:
-
match: The matched substring. (Corresponds to$0.) -
p1, p2, …, pN: The nth string found by a capture group (including named capturing groups) corresponds to$1,$2, etc. For example, if the pattern is/(\a+)(\b+)/, thenp1is the match for\a+, andp2is the match for\b+. If the group is part of a disjunction (e.g."abc".replace_with('/(a)|(b)/', replacer)), the unmatched alternative will benil. -
offset: The offset of the matched substring within the whole string being examined. For example, if the whole string was'abcd', and the matched substring was'bc', then this argument will be1. -
string: The whole string being examined.
The exact number of arguments depends on how many capture groups are contained in the regex.
For example:
%> echo 'name'.replace_with('/m/', @(match, offset) {
.. return match + '-'
.. })
'nam-e'
Below is another example that uses a capture group:
%> var text = 'all is well'
%>
%> echo text.replace_with('/([a-z]+)/', @(match, val) {
.. if val == 'is' return 'is not'
.. return 'will be'
.. })
'will be is not will be'
Parameters
regex(string) — The regular expression to match.callback(function) — The callback function to invoke for each match.
Returns string
ascii()
ascii() -> string
Reinterprets the string as a raw byte view: each byte of its UTF-8
encoding becomes its own character (a codepoint between 0 and 255,
i.e. a Latin-1-style one-byte-per-character mapping), rather than the
decoded sequence of Unicode characters that length(), each(), and
indexing otherwise operate on.
The result is still a valid string (every codepoint between 0 and
255 is valid UTF-8), so it can be used anywhere a normal string can.
It just no longer round-trips back through the original multi-byte
characters if the string had any, and its length() now reports the
original BYTE count of the string rather than its original CHARACTER
count.
This is meant for the rare case where code needs to walk a string byte-for-byte instead of character-by-character, e.g. one that originated from a byte stream where the bytes were never meant to be decoded as Unicode at all.
%> 'café'.length()
4
%> 'café'.ascii().length()
5
Returns string
case_fold()
case_fold() -> string
Returns a copy of the string case-folded for case-insensitive
comparison, using full Unicode case folding rather than plain
lowercasing. This matters for characters whose fold is not just their
lowercase form: for example, the German ß folds to ss.
Two strings that are considered equal ignoring case will always produce
identical output from case_fold(), which makes it the correct method
to use for case-insensitive comparisons; lower() is not a substitute
for it.
%> 'HELLO World'.case_fold()
'hello world'
%> 'Straße'.case_fold()
'strasse'
Returns string
compare()
compare(other: string) -> number
Compares the string with another string.
Parameters
other(string) — The other string to compare with.
Returns number — A negative number if the string is less than the
other string. - Zero if the strings are equal. - A positive number if
the string is greater than the other string.
Raises Error if the other string is not a string.
is_empty()
is_empty() -> boolean
Returns true if the string is empty, false otherwise.
Returns boolean
contains()
contains(str: string) -> boolean
Returns true if the string contains the specified substring, false otherwise.
Parameters
str(string) — The substring to search for.
Returns boolean
lines()
lines() -> list
Returns the lines of the string as an list as it would be if split on newline characters.
Returns list
each_line()
each_line(callback: function) -> void
Iterates over each line of the string, calling the provided callback function with the line and its index.
Parameters
callback(function) — A function that takes two arguments: the line and its index.
Returns void
Raises Error if the callback is not a function.
each()
each(callback: function) -> void
Iterates over each character of the string, calling the provided callback function with the character and its index.
Parameters
callback(function) — A function that takes two arguments: the character and its index.
Returns void
Raises Error if the callback is not a function.
capitalize()
capitalize() -> string
Returns a new string with the first character capitalized and the rest in lowercase.
Returns string
title()
title() -> string
Returns a new string with each word capitalized.
Returns string
to_string()
to_string() -> string
Returns the string itself.
Returns string