Strings
Text is bytes, and they are UTF-8
A string is immutable UTF-8 text with value semantics. A char is one Unicode scalar value.
let s = "héllo";
s.len() // 6: len counts bytes
s.char_at(0) // 'h': the char whose encoding holds byte 0
s[1] // 0xC3: the byte itself, a u8, for code that walks encodings
for c in s { // one iteration per character, decoded
println(`${c}`);
}A char literal holds one scalar value whatever its encoded width: 'é' as i64 is 233, and c.to_string() encodes a char back into the bytes it came from.
Literals and escapes:
let a = "line\nand \t tab, \"quoted\", \0 and \u{1F600}";
let c = '\n';Templates: the only formatting mechanism
let s = `${name} has ${count} items`;Each ${...} holds an expression of a primitive type or a type implementing Display; the pieces are concatenated into a string. There are no {} placeholders and no format!: this is the one mechanism in Core.
Interpolation is not a formatter: it answers six significant digits and switches to exponent notation on its own, so ${123456789.123} is 1.23457e+08. When the text matters, std.fmt says how many digits and in which base: fmt.float(v, 2) for fixed point, fmt.shortest(v) for the shortest text that reads back as the same double, fmt.base(v, radix) / fmt.hex, fmt.oct, fmt.bin, and fmt.group(text, 3, ",").
Operations
| Operation | Meaning |
|---|---|
s + t | concatenation |
s == t, s != t | content comparison |
s.len() | bytes |
s.char_at(i) | the char whose encoding holds byte i; '\0' past the end |
s[i] | the byte at i, a u8; past the end it panics, as xs[i] does. A u8 is not a char, so s[i] == 'a' is refused |
s.slice(a, b) | the text from byte a to byte b; a position past the end is the end, a negative one counts back from the end, and a at or after b is "" |
s.starts_with(t) | prefix test |
s.to_i64(), s.to_f64() | Option of the parsed number |
for c in s | iterate the characters, decoded |
A position inside a character means the start of that character, in char_at and in slice: "naïve".slice(0, 3) is "na" and "naïve".char_at(3) is 'ï'. Neither panics, and a slice is always UTF-8.
A string is reference counted by the runtime: a copy retains and a release decrements, which the program cannot observe because strings are immutable. A literal is a constant with an immortal count. .clone() is a deep copy.
Reading, building, and bytes that are not text
std.str.Cursor reads a string one character at a time, and its positions are always character starts: next() answers the character and moves past it, peek() answers it and stays, both None at the end; eat(ch) moves past ch when it is next; at() is the offset, and since(m) is the text from m to here. std.str.Builder makes a string by appending, with push(s), push_char(c) and to_string(): it is the answer to s = s + piece in a loop, which copies everything written so far on every piece.
std.bytes holds bytes that are not text: a buffer is an Array<u8>, bytes.of(s) copies a string's bytes, and bytes.text(buf, at, len) is the one way back to a string, an Err naming the offset when the bytes are not UTF-8. fs.read_file(p) reads a file as checked text, and fs.read(p) as bytes.
use std.str.{Cursor, Builder};
use std.bytes;
fn words(src: string): Array<string> {
let out: Array<string> = [];
let c = Cursor.of(src);
let going = true;
while going {
while c.eat(' ') { }
let start = c.at();
let more = true;
while more {
match c.peek() {
Some(ch) => { if ch == ' ' { more = false; } else { c.next(); } }
None => { more = false; going = false; }
}
}
if c.at() > start { out.push(c.since(start)); }
}
out
}
fn main(): i32 {
let b = Builder.new();
for w in words(" one sentence in words ") { b.push(w); b.push_char('|'); }
println(b.to_string());
let raw = bytes.of("ok");
raw.push(255);
match bytes.text(raw, 0, raw.len()) {
Ok(t) => println(t),
Err(e) => println(`not text: ${e.message()}`),
}
0
}