Skip to content

Strings ​

Text is bytes, and they are UTF-8 ​

A string is immutable UTF-8 text with value semantics. A char is one Unicode scalar value.

talor
let s = "héllo";
s.len()             // 6: len counts bytes
s.char_at(0)        // 'h': the char whose encoding holds byte 0
s[1]                // 0xC3: the byte itself, a u8, for code that walks encodings
for c in s {        // one iteration per character, decoded
    println(`${c}`);
}

A char literal holds one scalar value whatever its encoded width: 'é' as i64 is 233, and c.to_string() encodes a char back into the bytes it came from.

Literals and escapes:

talor
let a = "line\nand \t tab, \"quoted\", \0 and \u{1F600}";
let c = '\n';

Templates: the only formatting mechanism ​

talor
let s = `${name} has ${count} items`;

Each ${...} holds an expression of a primitive type or a type implementing Display; the pieces are concatenated into a string. There are no {} placeholders and no format!: this is the one mechanism in Core.

Interpolation is not a formatter: it answers six significant digits and switches to exponent notation on its own, so ${123456789.123} is 1.23457e+08. When the text matters, std.fmt says how many digits and in which base: fmt.float(v, 2) for fixed point, fmt.shortest(v) for the shortest text that reads back as the same double, fmt.base(v, radix) / fmt.hex, fmt.oct, fmt.bin, and fmt.group(text, 3, ",").

Operations ​

OperationMeaning
s + tconcatenation
s == t, s != tcontent comparison
s.len()bytes
s.char_at(i)the char whose encoding holds byte i; '\0' past the end
s[i]the byte at i, a u8; past the end it panics, as xs[i] does. A u8 is not a char, so s[i] == 'a' is refused
s.slice(a, b)the text from byte a to byte b; a position past the end is the end, a negative one counts back from the end, and a at or after b is ""
s.starts_with(t)prefix test
s.to_i64(), s.to_f64()Option of the parsed number
for c in siterate the characters, decoded

A position inside a character means the start of that character, in char_at and in slice: "naïve".slice(0, 3) is "na" and "naïve".char_at(3) is 'ï'. Neither panics, and a slice is always UTF-8.

A string is reference counted by the runtime: a copy retains and a release decrements, which the program cannot observe because strings are immutable. A literal is a constant with an immortal count. .clone() is a deep copy.

Reading, building, and bytes that are not text ​

std.str.Cursor reads a string one character at a time, and its positions are always character starts: next() answers the character and moves past it, peek() answers it and stays, both None at the end; eat(ch) moves past ch when it is next; at() is the offset, and since(m) is the text from m to here. std.str.Builder makes a string by appending, with push(s), push_char(c) and to_string(): it is the answer to s = s + piece in a loop, which copies everything written so far on every piece.

std.bytes holds bytes that are not text: a buffer is an Array<u8>, bytes.of(s) copies a string's bytes, and bytes.text(buf, at, len) is the one way back to a string, an Err naming the offset when the bytes are not UTF-8. fs.read_file(p) reads a file as checked text, and fs.read(p) as bytes.

talor
use std.str.{Cursor, Builder};
use std.bytes;

fn words(src: string): Array<string> {
    let out: Array<string> = [];
    let c = Cursor.of(src);
    let going = true;
    while going {
        while c.eat(' ') { }
        let start = c.at();
        let more = true;
        while more {
            match c.peek() {
                Some(ch) => { if ch == ' ' { more = false; } else { c.next(); } }
                None => { more = false; going = false; }
            }
        }
        if c.at() > start { out.push(c.since(start)); }
    }
    out
}

fn main(): i32 {
    let b = Builder.new();
    for w in words("  one sentence   in  words ") { b.push(w); b.push_char('|'); }
    println(b.to_string());

    let raw = bytes.of("ok");
    raw.push(255);
    match bytes.text(raw, 0, raw.len()) {
        Ok(t) => println(t),
        Err(e) => println(`not text: ${e.message()}`),
    }
    0
}

Talor v0.1.0 - Released under the MIT OR Apache-2.0 license.