Introduction
What BitterASM is
BitterASM is a metalanguage for constructing assembly languages. It defines the semantics and metaprogramming machinery needed to build an ISA, while assuming essentially nothing about the target architecture itself.
Traditional assemblers bake in a specific architecture: registers, instructions, opcodes, addressing modes, and binary encoding are all privileged, built-in concepts. That works fine until you want something the assembler’s authors didn’t anticipate — a pseudo-instruction, an alternate syntax, a non-binary target, a research architecture. Then you’re stuck extending the assembler itself.
BitterASM inverts this. The language core knows nothing about:
registers, instructions, opcodes, operands, immediates,
addresses, word sizes, endianness, calling conventions,
sections, labels, binary
Every one of those is defined in ordinary BitterASM code — as libraries, architecture packages, and evaluators. The complexity of a real ISA lives where it belongs: in the code describing that ISA, not in the language interpreting it.
The core ideas
Instructions are macros, not language features. mov isn’t a built-in instruction declaration — it’s a macro someone wrote and exposed publicly. There’s no fundamental distinction between a “real” instruction and a “pseudo” one at the language level; both are just macros that expand into some encoding.
Syntax is an interface, not part of the ISA. Because instructions are macros with their own invocation syntax, the same underlying x86 machinery could support mov rax, 3, an arrow-style rax <- 3, or something else entirely — all sharing one architecture implementation underneath.
Abstraction does not imply optimization. What you expand is what you get. If a macro expands clear rax into xor rax, rax, that’s the macro author’s explicit choice. BitterASM never second-guesses it by picking a “faster” sequence on its own. A macro can implement sophisticated code selection, but that logic belongs to the macro author, not the language.
Even bits are a library concept. The only thing BitterASM assumes is Int — an architecture-neutral, arbitrary-precision integer. Binary (bit, bits<N>) is defined in terms of Int in the standard library, not built into the language. 3 is just the mathematical integer three until some library gives it a binary interpretation. This leaves room for architectures that aren’t binary at all.
Evaluators own the output contract, not the language. BitterASM source describes behavior; an evaluator decides what running that behavior produces. A binary evaluator turns assembly into machine code. A different evaluator could target something else entirely. Language semantics (Int, struct, type, macro, pattern matching, imports, generics) mean the same thing everywhere — only the output effect changes.
Imports compose modules; they don’t grant magic. from x86_64.intel import * works because that package defines mov and rax — not because the language has special knowledge of x86. Importing a binary library doesn’t secretly switch the language into “binary mode”; binary is just one ordinary abstraction among many.
Why bother
The payoff for assuming almost nothing is:
- Portability — the same metaprogramming core builds tiny pedagogical ISAs, real architectures like RISC-V, and pathologically complex ones like x86-64, without special-casing any of them.
- Auditable, structured assembly — instructions and their expansions are ordinary, inspectable code rather than opaque tables inside an assembler binary.
- Alternate syntaxes for free — because syntax lives with the macros that define it, an architecture can offer multiple front-ends (native, AT&T-style, “pretty,” even something Pythonic) over one shared implementation.
- Room for the unknown — future architectures that aren’t neatly binary, register-based, or von Neumann at all don’t require changes to BitterASM itself.
The tradeoff is symmetric: the fewer assumptions the language makes, the more an architecture package has to define for itself. That’s the bitter part. The resulting portability, auditability, and extensibility are the sweet part.
Getting started
BitterASM comes as two programs:
bitterasm, the compiler. It reads.basmsource, expands every macro, and writes the values the program emits to a.emfile.bitter, an evaluator. It reads a.emfile and packs its values into machine-code bytes.
Keeping them separate is deliberate. bitterasm knows nothing about bits,
bytes or machine code: it only produces a stream of values. What those values
mean is up to whichever evaluator reads them, and bitter is the one that
turns them into binary. See Evaluators.
This chapter covers:
- Installation: building and installing both programs and the standard library.
- Your first program: assembling two RISC-V instructions, and then a program that emits plain numbers.
- The toolchain: every
bitterasmandbittercommand.
Installation
BitterASM is built from source and needs a Rust toolchain (cargo). From a
clone of the repository:
make install # or: ./install.sh
The script asks before installing each part:
bitterasm(the compiler) andbitter(the binary evaluator), into~/.bitterasm/bin.- The standard library, copied from
std/to~/.bitterasm/std. - Optionally, the
bitterasm-lsplanguage server.
Pass -y to answer yes to every prompt. Add ~/.bitterasm/bin to your
PATH if the script tells you to.
Where imports are found
An absolute import such as from std.riscv.native import * is looked up in
these directories, in order:
- The current directory.
- Each directory in
BITTERASM_PATH(separated likePATH). ~/.bitterasm.
So a program anywhere on disk finds the installed std, and a project’s own
std/ directory takes priority over it. See Imports.
Checking it works
bitterasm --version
bitter --version
Next: Your first program.
Your first program
Assembling real instructions
Save this as first.basm:
from std.riscv.native import *
add a0, a1, a2
addi a0, a0, 1
33 85 c5 00 13 05 15 00
Then compile it and pack the result:
bitterasm compile first.basm # writes first.em
bitter encode first.em # writes first.bin
first.bin holds the eight bytes above: two RV32I instructions,
little-endian, exactly what a RISC-V assembler would produce.
Nothing in the language itself knows what add or a0 is. Both are
ordinary declarations in the standard library: add is a
macro and a0 is a constant.
The import brings them into scope, like importing a library in any other
language.
Emitting plain values
A program’s output is whatever its macros emit. Here is a program with no architecture at all:
macro show(value: int) {
@emit value
}
show 1
show 2 + 3
show 'A'
1 5 65
macro show(value: int) { ... }declares a macro namedshowthat takes one integer.@emitadds a value to the program’s output.show 1calls the macro. A line that starts with a macro’s name followed by its arguments is a call, just like an instruction in a traditional assembler.
Compile it and look at the .em file:
bitterasm compile show.basm
cat show.em
{
"version": 1,
"requires": [],
"module": "show",
"exports": {},
"entries": [
{ "kind": "Int", "value": "1" },
{ "kind": "Int", "value": "5" },
{ "kind": "Int", "value": "65" }
]
}
bitter encode show.em fails, and that’s on purpose. A bare integer has no
width, so bitter can’t tell how many bits 5 should take up. Giving it a
width is a library’s job: std.binary defines bits<N>, and the RISC-V
package builds its instructions out of bits<N> fields. See
Packing bytes with bitter.
Next: The toolchain.
The toolchain
bitterasm
| Command | What it does |
|---|---|
bitterasm compile prog.basm [-o prog.em] | Expands the program and writes its emitted values to a .em file. |
bitterasm check prog.basm | Runs every check compile does, but writes nothing. |
bitterasm expand prog.basm | Prints the program with every macro call replaced by the macro’s body. Nothing is evaluated. |
bitterasm doc <paths> [-o dir] | Writes reference pages from doc comments; --test compiles their examples. See Generating docs. |
bitterasm format <paths> | Formats .basm files in place (fmt for short). See Formatting. |
compile and check also take lint options (-A, -W, -D, -F) and
output options (--diagnostic-format, --color). See
Diagnostics and lints. compile --verbose shows
progress and timing for each top-level call.
bitter
| Command | What it does |
|---|---|
bitter encode prog.em [-o prog.bin] | Packs one .em file’s values into bytes. |
bitter build a.basm [b.basm ...] [-o prog] | Compiles every input, links them, packs the result and marks it executable. |
bitter build runs the whole pipeline, so you don’t need the intermediate
.em files. With several inputs, it joins their same-named
sections and resolves pub
labels across files. See
Linking multiple files and
Executables.
A typical session
bitterasm check hello.basm # fast feedback while editing
bitter build hello.basm -o hello # compile, link and pack
./hello
Basics
A BitterASM program is a text file, usually ending in .basm. This chapter
covers what every program is made of: its layout, and the values it
computes with.
One statement per line
A newline ends a statement, just as in a traditional assembler:
macro show(value: int) {
@emit value
}
show 1
show 2
1 2
Inside (), [] and {}, a statement can continue onto the next line, so a
long call can be split:
macro add3(a: int, b: int, c: int) -> int {
@return a + b + c
}
macro show(value: int) {
@emit value
}
show add3(
1,
2,
3
)
6
Comments
# starts a comment that runs to the end of the line. There are no block
comments. ## and #! start doc comments.
# A whole-line comment.
macro show(value: int) {
@emit value # a trailing comment
}
show 7
7
Names
Names may contain letters, digits and underscores, and don’t start with a
digit. Any Unicode letter counts, so π is a valid name. A name that starts
with an underscore marks a parameter as deliberately unused (see
Diagnostics and lints).
The only reserved words are from, import, as, in, pub, skip,
macro, type, struct, enum, const and section. Everything else,
including int and every instruction mnemonic, is an ordinary name.
Order doesn’t matter
A file’s declarations are visible throughout the file, so you can use a macro or constant above the line that declares it:
show LIMIT
const LIMIT = 9
macro show(value: int) {
@emit value
}
9
What does follow source order is the output: calls emit their values in the order they appear.
What a file can contain
| Statement | Example | Chapter |
|---|---|---|
| Import | from std.binary import bits | Modules |
| Constant | const WIDTH = 8 | Constants |
| Macro | macro nop() { ... } | Macros |
| Struct, enum, type alias | struct Point { ... } | Types |
| Macro call | add a0, a1, a2 | Macros |
| Label | loop: | Labels |
| Section | section .text | Sections |
| Syntax override | syntax add(a, b) = { ... } | Custom syntax |
@for, @if, @fold | @for i in 0..4 { ... } | Meta keywords |
In this chapter
- Integers: the one built-in type.
- Operators: arithmetic, bitwise, comparison and logic.
- Ranges and
in:0..n, and testing membership. - Characters and strings.
- Constants: naming values with
const.
Integers
int is BitterASM’s only built-in type. An int is a mathematical integer:
it has arbitrary precision and no width, so it never overflows or wraps.
macro show(value: int) {
@emit value
}
show 1_000_000 * 1_000_000 * 1_000_000 * 1_000_000
1000000000000000000000000
3 is just the number three, with no bits attached. Bits, bytes, signed and
unsigned values are all library types built on int: see bits<N> in
Packing bytes with bitter, and
Types for how to build types of your own.
Literals
| Form | Example | Value |
|---|---|---|
| Decimal | 42 | 42 |
| Hexadecimal | 0xff | 255 |
| Binary | 0b1010 | 10 |
| Octal | 0o17 | 15 |
| Character | 'A' | 65 |
Any literal may use _ between digits for readability: 1_000_000,
0xffff_0000, 0b1010_0101.
macro show(value: int) {
@emit value
}
show 42
show 0xff
show 0b1010
show 0o17
show 0xff_ff
42 255 10 15 65535
Negative numbers
There are no negative literals. -5 is the negation operator applied to 5,
which gives the same result:
macro show(value: int) {
@emit value
}
show -5
show -(2 + 3)
-5 -5
True and false
There’s no built-in boolean. Conditions are integers: 0 is false and
anything else is true. Comparisons and logical operators produce 0 or 1.
See Operators.
std.binary does define a bool type (a one-bit bits<1>) with true and
false constants, for code that needs a boolean it can emit.
No floating point
There are no fractional numbers in the language. std.decimal provides
Decimal and Fraction types, and std.math provides fixed-point
arithmetic, square roots, logarithms and trigonometry over them.
Operators
Arithmetic
| Operator | Meaning | Example | Result |
|---|---|---|---|
+ | Addition | 7 + 2 | 9 |
- | Subtraction | 7 - 2 | 5 |
* | Multiplication | 7 * 2 | 14 |
/ | Division, rounding toward zero | -7 / 2 | -3 |
% | Remainder, with the sign of the left side | -7 % 2 | -1 |
-x | Negation | -7 | -7 |
macro show(value: int) {
@emit value
}
show 7 / 2
show -7 / 2
show -7 % 2
3 -3 -1
For floor, ceiling or Euclidean division, use std.math.
Bitwise
Integers have no width, so bitwise operators act as if every number had infinitely many bits, in two’s complement. Negative numbers have infinitely many leading ones.
| Operator | Meaning | Example | Result |
|---|---|---|---|
& | And | 0b1100 & 0b1010 | 8 |
| | Or | 0b1100 | 0b1010 | 14 |
^ | Exclusive or | 0b1100 ^ 0b1010 | 6 |
~x | Not | ~0 | -1 |
<< | Shift left | 1 << 4 | 16 |
>> | Shift right, keeping the sign | -16 >> 2 | -4 |
macro show(value: int) {
@emit value
}
show 0b1100 & 0b1010
show 0b1100 | 0b1010
show 0b1100 ^ 0b1010
show ~0
show 1 << 4
show -16 >> 2
8 14 6 -1 16 -4
Comparison
==, !=, <, <=, > and >= produce 1 for true and 0 for false.
== and != also compare structs and enums, field by field.
macro show(value: int) {
@emit value
}
show 2 < 3
show 2 == 3
1 0
Logic
| Operator | Meaning |
|---|---|
&& | 1 if both sides are true, else 0 |
|| | 1 if either side is true, else 0 |
!x | 1 if x is 0, else 0 |
Any nonzero value counts as true, and the result is always 0 or 1:
macro show(value: int) {
@emit value
}
show 3 && 4
show 0 || 7
show !5
1 1 0
Precedence
From tightest to loosest binding:
| Level | Operators |
|---|---|
| Postfix | .field, calls f(x), as |
| Prefix | -x, !x, ~x |
| Multiplicative | *, /, % |
| Additive | +, - |
| Shift | <<, >> |
| Comparison | <, <=, >, >=, in |
| Equality | ==, != |
| Bitwise and | & |
| Bitwise xor | ^ |
| Bitwise or | | |
| Logical and | && |
| Logical or | || |
| Range | .., ..= |
Operators on the same level group left to right. Use parentheses whenever the grouping isn’t obvious, especially when mixing bitwise and comparison operators:
macro show(value: int) {
@emit value
}
show 1 << 4 | 1
show (6 & 3) == 2
17 1
Since as binds tightly, a + b as T means a + (b as T), and
-1 as T means -(1 as T). Write (a + b) as T and (-1) as T. See
Conversions.
from std.ctypes import int8_t
macro emit_byte(b: int8_t) {
@emit b
}
emit_byte -1 as int8_t
`as` binds more tightly than `-`, so `-x as T` means `-(x as T)`; write `(-x) as T`
Ranges and in
Ranges
A range is a run of consecutive integers:
| Syntax | Contains |
|---|---|
a..b | a, a + 1, …, b - 1 (excludes b) |
a..=b | a, a + 1, …, b (includes b) |
If b is not greater than a (or, for ..=, less than a), the range is
empty. Ranges only count upward. For other steps, use std.iter’s range.
A range is mostly used as the source of a @for loop:
macro show(value: int) {
@emit value
}
@for i in 0..3 {
show i
}
@for i in 0..=3 {
show i * 10
}
0 1 2 0 10 20 30
Inside a macro, a range can be stored in a constant and used later. A
top-level @for needs its range written in place, because it’s unrolled
before anything else in the file is evaluated.
in
value in source is 1 if source contains value, and 0 otherwise.
The source can be:
- A range.
i in 0..lenchecks0 <= i && i < len. Nothing is iterated, so it’s cheap even for a range with a trillion elements. - A struct value, such as an array or a string.
x in arris true ifxequals one of the elements@forwould visit: thepubfields that aren’tskip, as described in Struct fields.
macro show(value: int) {
@emit value
}
show 5 in 0..5
show 5 in 0..=5
show 1000 in 0..1_000_000_000_000
show 'b' in "abc"
show 'z' in "abc"
0 1 1 1 0
in is often used in an @assert:
macro checked_index(idx: int, len: int) {
@assert idx in 0..len, "index out of bounds"
@emit idx
}
checked_index 3, 4
3
macro checked_index(idx: int, len: int) {
@assert idx in 0..len, "index out of bounds"
@emit idx
}
checked_index 4, 4
index out of bounds
in binds as tightly as <, and a range binds more loosely than anything
else, so i + 1 in 0..n + 1 means (i + 1) in 0..(n + 1).
in isn’t for types
x in SomeEnum doesn’t test whether x is a valid variant. in always asks
whether a value you have contains something, the same question @for
walks through. Asking whether a value belongs to a type is a different
question, so it doesn’t share the keyword.
Characters and strings
Characters
A character literal like 'a' is just an int: the character’s Unicode
code point. '€' is 0x20AC, whatever machine you target. Characters carry
no encoding and no byte order.
macro show(value: int) {
@emit value
}
show 'A'
show '€'
show '\n'
65 8364 10
A character literal holds exactly one character. The escapes are \n, \r,
\t, \0, \\ and \'.
Strings
A string literal like "abc" is a struct with one int field per character
(its code point) and a len field:
| Field | Value for "hé€" |
|---|---|
__el0 | 'h' (104) |
__el1 | 'é' (233) |
__el2 | '€' (8364) |
len | 3 |
len counts characters, not bytes. Strings accept the same escapes as
characters, with \" in place of \'.
macro show(value: int) {
@emit value
}
const text = "hé€"
show text.len
show text.__el1
3 233
@for visits a string’s characters, but not its len,
which is declared pub skip (see Struct fields):
macro each_char() {
@for c in "ab" {
@emit c
}
}
each_char
97 98
Strings as bytes
A string literal has no encoding until a library gives it one:
std.array’sarray_from_struct("abc")turns it into anArray<int, 3>.std.string’sstring_from_struct("abc")encodes it as UTF-8, packed into one integer, with alenin bytes.
from std.string import string_from_struct
macro packed() {
const s = string_from_struct("hi")
@emit s.value
@emit s.len
}
packed
26729 2
26729 is 0x6869: the bytes h (0x68) and i (0x69).
To put a string’s bytes in a program, an architecture package provides a
directive for it. For example, std.x86_64.nasm has NASM’s db:
msg:
db "Hello, World!\n"
Constants
const gives a value a name:
const WIDTH = 8
const MASK = (1 << WIDTH) - 1
macro show(value: int) {
@emit value
}
show MASK
255
A constant can hold any value: an integer, a struct, an enum variant, or a string.
Syntax
const NAME = value
const NAME: Type = value
pub const NAME = value
: Typeconverts the value toType, and checks it, before binding it. See Conversions.publets other files import the constant. See Visibility.
from std.binary import bits
const BYTE: bits<8> = 200
macro emit_byte(b: bits<8>) {
@emit b
}
emit_byte BYTE
bits<8> { value: 200 }
A value that doesn’t fit the type is a compile error. bits<8> checks that
its value fits in eight bits:
from std.binary import bits
const BYTE: bits<8> = 300
invariant `fits_inside_width(width, value)` was violated for `bits`
Constants never change
There are no variables in BitterASM. A constant is bound once and never
reassigned, and nothing in the language can be mutated. To compute a value
step by step, use recursion or @fold.
Constants inside macros
Inside a macro body, const names an intermediate value. It’s visible for
the rest of that body:
macro hypot_squared(a: int, b: int) {
const aa = a * a
const bb = b * b
@emit aa + bb
}
hypot_squared 3, 4
25
A pub const inside a macro is different: it declares a new top-level
constant when the macro is called. See
Generating declarations.
Registers are constants
Architecture packages use constants for registers. For example,
std.riscv.impl declares each register as a constant of type Reg, a
five-bit number:
pub type Reg = bits<5>
@for i in 0..32 {
pub const x`i` = Reg(i)
}
pub const zero = x0
pub const ra = x1
The x`i` builds the names x0 to x31. See
Spliced names.
Doc comments
## documents the item below it, and #! documents the whole file. The
text is Markdown.
Syntax
#! What this file is for.
## What this item does.
item
Example
#! Helpers for showing values.
## Emits `value` unchanged.
##
## ```
## show 7
## ```
macro show(value: int) {
@emit value
}
show 7
7
Doc comments are for the people who use your code. Plain # comments are
for the people who maintain it, and are never part of the docs. The two mix
freely:
## Emits `value` twice.
# Written as two `@emit`s rather than a loop, to keep the expansion short.
macro twice(value: int) {
@emit value
@emit value
}
twice 3
3 3
What can be documented
A ## block documents the next item, as long as only blank lines and #
comments sit in between. An item is any of:
- a
macro,struct,enum,typealias orconst - a label
- a
syntaxline - a struct field or an enum variant
## A 2D point.
struct Point {
## Distance from the left edge.
pub x: int,
## Distance from the top edge.
pub y: int,
}
## Where a program starts.
enum Entry {
## At the first instruction.
Start,
## At an offset from it.
Offset: int,
}
macro show(value: int) {
@emit value
}
show Point { x: 1, y: 2 }.y
2
An item inside a block is documented the same way, such as each constant
a top-level @for generates:
@for i in 0..4 {
## Register `i`.
pub const r`i` = i
}
macro show(value: int) {
@emit value
}
show r3
3
#! lines only document the file when they come before its first
statement.
To turn doc comments into reference pages, see Generating docs.
Where doc comments are ignored
A doc comment that doesn’t document anything is reported by the
unused_doc_comments lint. That happens when:
- a
##sits above something that isn’t an item, such as an invocation or an@emit - a
##is at the end of a file or a block - a
#!comes after the file’s first statement
macro show(value: int) {
@emit value
}
## Shows seven.
show 7
doc comment documents nothing
Use # for an ordinary comment.
### and longer runs of # are ordinary comments too, so banners such as
#### Registers #### never end up in the docs.
Formatting
bitterasm fmt keeps each line’s ## or #! marker. It wraps prose that
runs past comment_width, but leaves code blocks, tables and headings
exactly as written.
Macros
Macros are how BitterASM does everything. An instruction like add a0, a1, a2
is a call to a macro named add. A “pseudo-instruction” is a macro too,
and so is a helper that computes a value. The language draws no line between
them.
A macro runs at compile time. It can compute values, check conditions, call other macros, and emit values into the program’s output.
Declaring a macro
macro name(param: Type, ...) -> ReturnType
| facet ...
{
body
}
- Parameters each have a name and a type. See Parameters and defaults.
-> ReturnTypeis optional. See Returning values.- Facets (
| before ...,| syntax { ... }, …) are optional modifiers, one per line, before the body. pub macrolets other files import it. See Visibility.
Two ways to call a macro
As a statement, like an instruction: the name, then its arguments separated by commas, with no parentheses. Whatever the macro emits goes into the output.
As an expression, like a function: the name, then its arguments in parentheses. The call’s value is whatever the macro returns.
macro square(x: int) -> int {
@return x * x
}
macro show(value: int) {
@emit value
}
show 3 # statement call
show square(4) # `square(4)` is an expression call
3 16
A statement call throws away the macro’s return value, so a statement call
of square would do nothing. An expression call can’t throw away emitted
values, so a macro that emits can only be called as an expression in a few
places. See Where emitted values can go.
A macro with no parameters is called by its name alone:
macro nop() {
@emit 0x13
}
nop
nop
19 19
Inside a macro
A macro’s body is a list of statements:
- Meta keywords such as
@emit,@return,@ifand@for. - Calls to other macros, which emit into this macro’s output.
constdeclarations naming intermediate values.- Declarations that the macro generates when it’s called. See Generating declarations.
macro show(value: int) {
@emit value
}
macro countdown(start: int) {
@for i in 0..start {
show start - i
}
show 0
}
countdown 3
3 2 1 0
In this chapter
- Parameters and defaults
- Returning values
- Overloading: several macros with one name.
- Recursion: macros that call themselves.
- Hooks:
beforeandafter: code that runs around a macro. - Custom syntax: calls shaped like
mov rax, [rbx]orx <- 5. - Generating declarations: macros that declare constants and types.
Macros can also be generic, and can take other macros as arguments. See Generics.
Parameters and defaults
Parameters
Every parameter has a name and a type:
macro show_sum(a: int, b: int) {
@emit a + b
}
show_sum 2, 3
5
Arguments are matched to parameters by position. Named arguments like
f(b = 1) are only for constructing structs, not for
calling macros.
Arguments are type-checked
Each argument must already have its parameter’s type. Nothing is converted automatically, even when a conversion exists:
from std.binary import bits
macro emit_byte(b: bits<8>) {
@emit b
}
emit_byte 65
type mismatch for `b`: expected `bits<8>`, found `int`
Convert explicitly with as:
from std.binary import bits
macro emit_byte(b: bits<8>) {
@emit b
}
emit_byte 65 as bits<8>
bits<8> { value: 65 }
Defaults
Parameters at the end of the list can have a default value, used when the call leaves them out:
macro encode(value: int, width: int = 8, mask: int = (1 << width) - 1) {
@emit value & mask
}
encode 0x1ff
encode 0x1ff, 4
255 15
- A default can use earlier parameters. Defaults are evaluated left to right
at each call, so
maskabove sees whicheverwidththe call ended up with. - Once one parameter has a default, every parameter after it needs one too.
Unused parameters
An unused parameter triggers the unused_parameter warning. If it’s unused
on purpose, for example because only its type matters for
overloading, start its name with an underscore:
macro kind(_value: int) {
@emit 1
}
kind 42
1
Returning values
A macro gives a value back to its caller with @return:
macro max(a: int, b: int) -> int {
@if a > b {
@return a
}
@return b
}
macro show(value: int) {
@emit value
}
show max(3, 9)
9
-> int declares the return type. It’s optional, but if it’s there, it’s
enforced. It’s also needed when the macro is passed as an argument (see
Macros as parameters) or used as a
conversion.
The returned value must have the declared type. As with arguments, nothing is converted automatically:
from std.binary import bits
macro byte() -> bits<8> {
@return 5
}
macro emit_byte(b: bits<8>) {
@emit b
}
emit_byte byte()
`byte` returned `int`, but its signature declares `-> bits<8>`
Write @return 5 as bits<8> instead.
-> is only about returning. A macro that declares -> T must return a
T, so declaring a return type on a macro that only emits is an error:
from std.binary import bits
macro nop() -> bits<8> {
@emit 0x90 as bits<8>
}
nop
`nop` returned nothing, but its signature declares `-> bits<8>`
To declare what a macro emits, use the
emits facet
instead: macro nop() | emits bits<8> { ... }.
Returning versus emitting
Every macro has two separate outputs:
@return | @emit | |
|---|---|---|
| Goes to | The caller, as the call’s value | The program’s output |
| How many | One value, or none | Any number |
| Ends the macro | Yes | No |
A macro may do both. A macro that only emits is like an instruction; a macro that only returns is like a function.
Where emitted values can go
A statement call puts the macro’s emitted values in the output where the call appears. A call used as an expression has nowhere to put emitted values unless it’s the whole of one of these:
- a
const’s value:const x = f() @return’s value:@return f()
In both cases, the emitted values are kept, in order, where that statement
appears. The same goes for @fold in expression position.
macro emit_and_return(x: int) -> int {
@emit x
@return x * 10
}
macro caller() {
const y = emit_and_return(1)
@emit y
}
caller
1 10
Anywhere else, such as inside a larger expression, calling a macro that emits is an error, because its values would be lost:
macro emit_and_return(x: int) -> int {
@emit x
@return x * 10
}
macro caller() {
@emit 1 + emit_and_return(1)
}
caller
`emit_and_return` emits values, so it can only be used as a statement
Macros that return nothing
Using a macro that returns nothing as a value is an error. A bare @return
ends a macro early without a value.
Overloading
Several macros can share a name, as long as their parameters differ. A call picks the one whose parameter types and count fit its arguments:
struct Reg {
pub n: int,
}
macro describe(r: Reg) {
@emit 1000 + r.n
}
macro describe(v: int) {
@emit v
}
describe Reg(3)
describe 3
1003 3
This is how instruction sets handle operand forms: std.x86_64 has one mov
for register-to-register, another for an immediate, another for memory, and
so on, each with its own encoding.
How a call picks an overload
- Only overloads that accept the number of arguments are considered, counting defaults.
- Of those, only overloads whose parameter types match the arguments are considered.
- A non-generic overload beats a generic one.
- If more than one candidate is left, the call is ambiguous, which is an error.
macro kind<T>(_x: T) {
@emit 1
}
macro kind(_x: int) {
@emit 2
}
kind 5
2
macro g(_a: int, _b: int = 0) {
@emit 3
}
macro g(_a: int) {
@emit 4
}
g 1
multiple overloads of `g` accept (int)
Overloads across files
A file’s overloads merge with same-named overloads it imports, from any
number of modules. That’s how a dialect adds new operand forms to an
instruction it builds on: std.x86_64.nasm adds mov rax, [rel label]
alongside every mov from std.x86_64.intel. See
Imports.
Recursion
A macro can call itself, directly or through other macros:
macro factorial(n: int) -> int {
@if n <= 1 {
@return 1
}
@return n * factorial(n - 1)
}
macro show(value: int) {
@emit value
}
show factorial(20)
2432902008176640000
Limits
Since macros run inside the compiler, a runaway recursion must not crash it, so calls are limited:
- At most 32 nested calls. Deeper than that fails with
MacroCallDepthExceeded. The limit counts every nested call, whether it was made as a statement or as an expression. - Tail calls don’t count. A macro whose
@returnis exactly a call to itself, like@return gcd(b, a % b), reuses the current call instead of nesting a new one. A tail-call loop can run up to 4,096 times before it fails withMacroTailCallLimitExceeded.
macro gcd(a: int, b: int) -> int {
@if b == 0 {
@return a
}
@return gcd(b, a % b)
}
macro show(value: int) {
@emit value
}
show gcd(1071, 462)
21
A call is only a tail call if it’s the whole @return value.
@return 1 + f(x) isn’t one. A macro with an after hook never
makes tail calls, because its hooks have to run after each call finishes.
Prefer loops for long runs
For anything that repeats more than a few thousand times, use
@for or @fold instead. They have no
depth limit, and can run up to 1,000,000 iterations:
macro sum_to(n: int) -> int {
@return @fold total = 0 @for i in 0..=n {
@next total + i
}
}
macro show(value: int) {
@emit value
}
show sum_to(100_000)
5000050000
Hooks: before and after
A macro can declare calls that run every time it’s called: before hooks
run first, then the body, then after hooks.
macro show(value: int) {
@emit value
}
macro traced(x: int) -> int
| before show(100)
| after show(result.returned)
{
@return x + 1
}
traced 1
100 2
Syntax
macro name(params)
| before hook_call(params...)
| after hook_call(params..., result.returned)
{
body
}
- A hook is a macro call. It can use the macro’s parameters.
- In an
afterhook,result.returnedis the value the body returned. - A macro can have any number of each. They run in the order they’re written.
- Anything a hook emits goes into the output, like the body’s own emits.
Checking arguments
The most common use is a shared check. std.array checks every index once,
in a hook, rather than in each macro’s body:
macro oob_check(len: int, index: int) {
@assert index >= 0, "negative index"
@assert index < len, "index past the end"
}
macro element_offset(len: int, index: int) -> int
| before oob_check(len, index)
{
@return index * 4
}
macro show(value: int) {
@emit value
}
show element_offset(8, 3)
12
Hooks compose
A hook is an ordinary macro call, so if the hook macro has hooks of its own, they run too. Checks built from hooks keep working however deep the call chain goes.
Hooks change when a macro can tail-call itself: a
macro with an after hook never does, because the hook has to run after
each call returns.
Custom syntax
By default, a macro is called as its name followed by comma-separated
arguments: swap 1, 2. The syntax facet gives a macro any call shape you
like. That’s how assembly syntax is built: lw a0, 8(sp),
mov rax, [rbx + 8] and x1 = x2 + x3 are all ordinary macros with custom
syntax.
macro swap(a: int, b: int)
| syntax { swap $a$ with $b$ }
{
@emit b
@emit a
}
swap 1 with 2
2 1
Patterns
A pattern is a sequence of tokens between { and }:
$name$is a capture. The call site can put any expression there, and it becomes the argument for parametername.- Anything else is literal: the call site must contain exactly that token.
Every capture must name one of the macro’s parameters. To match a literal
` or $, escape it: \` or \$.
Three kinds of pattern
Anchored patterns start with the macro’s own name, like the swap
pattern above. They’re tried only for statements that start with that name,
so they’re cheap. Most instruction syntax is anchored.
Unanchored patterns start with something else, usually a capture. They are tried against every statement, so a statement doesn’t need to start with a mnemonic at all:
macro store(dst: int, src: int)
| syntax { $dst$ <- $src$ }
{
@emit dst * 100 + src
}
const r1 = 1
r1 <- 7
107
A statement must start with a name, so r1 <- 7 works where 1 <- 7
wouldn’t.
Operand patterns start with a token that can’t begin a statement, such
as [. They’re tried wherever an argument starts, and a match becomes an
expression call. So an operand pattern belongs on a macro that returns a
value:
struct Mem {
pub addr: int,
}
macro mem(addr: int) -> Mem
| syntax { [$addr$] }
{
@return Mem(addr)
}
macro load(dst: int, src: Mem) {
@emit dst
@emit src.addr
}
load 1, [0x40]
1 64
[0x40] becomes mem(0x40), a Mem, and overloading
picks the load that takes a Mem. This is how std.x86_64.nasm gives
[rel label] its own type.
Changing another macro’s syntax
A syntax statement assigns a call shape to a macro declared somewhere
else, without touching its declaration:
macro copy(dst: int, src: int) {
@emit dst
@emit src
}
syntax copy(dst, src) = { copy $src$ to $dst$ }
copy 5 to 6
6 5
The names in parentheses are the macro’s parameters, which the pattern can
capture. This is what makes dialects possible. std.riscv.impl declares
every instruction with no syntax of its own. Then:
std.riscv.nativeassigns conventional syntax:lw a0, 8(sp).std.riscv.c_likeassigns C-like syntax to the same macros:a0 = a1 + a2.
A program imports one dialect or the other, and both share one
implementation. A file that imports a macro gets the syntax its dialect
assigned too. If two imports assign different syntax to the same macro,
the importing file must pick one with a syntax statement of its own.
Designing patterns
A capture accepts any expression, and it stops only where the next literal token appears. Two patterns are ambiguous when some statement matches both, so start them differently:
std.riscv.c_likewrites register forms as$rd$ = $rs1$ + $rs2$and immediate forms as$rd$ <- $rs1$ + $imm$. If both used=,x1 = x2 + imm - 1would be a valid parse of either one.- Avoid a literal
:right after a leading capture.name:is how a label is written, and that check runs first.
Declare a macro with custom syntax before any statement that uses it in the same file. The parser can misread an earlier call that uses a shape it hasn’t seen yet.
Generating declarations
A macro body can contain declarations: pub const, struct, enum,
type, and even macro. Each call to the macro adds them to the program as
if they had been written at the top level.
macro declare_word(bits: int) {
pub const WORD`bits`_BYTES = `bits / 8`
struct Word`bits` {
pub value: int,
}
}
declare_word 16
declare_word 32
macro show(value: int) {
@emit value
}
macro emit_word(w: Word16) {
@emit w
}
show WORD32_BYTES
emit_word Word16(7)
4
Word16 { value: 7 }
Two calls generated WORD16_BYTES, Word16, WORD32_BYTES and Word32.
What gets evaluated
When the macro runs, the declaration’s name is evaluated: the backticks
in Word`bits` splice in the parameter’s value. See
Spliced names.
Everything else is copied into the program as written, and evaluated later where the declaration ends up. There, the macro’s parameters no longer exist.
The one exception is a pub const’s value. Backticks in it are evaluated
when the macro runs, so `bits / 8` above becomes 2 or 4. Without
the backticks, bits / 8 would refer to a bits that isn’t defined at the
top level.
A generated struct, enum, type alias or macro gets no such exception: only its name can use the macro’s parameters.
Inside a macro, pub const and const differ
const x = ...names a value for the rest of the macro body.pub const x = ...generates a top-level constant.
Seeing what was generated
A program that generates declarations gets a generated_declarations
warning, as a reminder that it contains code you can’t see in its source.
bitterasm expand prints the program with each call replaced by its body,
which shows what was generated:
bitterasm expand program.basm
To silence the warning, allow the lint (see Diagnostics and lints).
Meta keywords
Words starting with @ are meta keywords. They control what a macro
does while the compiler runs it: emit values, return, check conditions,
branch and loop. There are eight:
| Keyword | Does | Page |
|---|---|---|
@emit value | Adds value to the program’s output | @emit |
@return [value] | Ends the macro, optionally with a value | @return |
@assert cond[, "message"] | Fails compilation if cond is false | @assert |
@if cond { } @else { } | Runs one branch | @if and @else |
@match value { pattern => { } } | Runs the first arm that matches | @match |
@for x in source { } | Runs the body once per element | @for |
@fold acc = init @for x in source { } | A @for that carries values between iterations | @fold and @next |
@next value | Moves a @fold to its next iteration | @fold and @next |
Nothing a meta keyword does survives into the output except what it emits.
There’s no @if at run time on the target machine; that would be an
instruction, which is a macro some architecture package provides.
Where they can be used
| Macro body | Top level | Struct declaration | Struct construction | |
|---|---|---|---|---|
@emit, @return, @assert | ✓ | |||
@if, @for, @fold/@next | ✓ | ✓ | ✓ | ✓ |
@match | ✓ | ✓ |
@emit, @return and @assert only mean something while a macro runs:
const WIDTH = 8
@assert WIDTH % 8 == 0
`@assert` can only be used inside a macro
At the top level, @for and @if repeat or choose statements: they can
generate declarations and calls. In a struct declaration they choose
fields, and in a construction they choose field values. See
Structs.
macro show(value: int) {
@emit value
}
@for i in 0..3 {
show i * i
}
@if 2 > 1 {
show 100
}
0 1 4 100
@emit
@emit adds one value to the program’s output.
Syntax
@emit value
Example
macro bytes3(a: int, b: int, c: int) {
@emit a
@emit b
@emit c
}
bytes3 1, 2, 3
1 2 3
What can be emitted
Any value: an integer, a struct, an enum variant. The compiler doesn’t care
what a value means; it just records it in order in the .em file. What the
value turns into is the evaluator’s decision. For
example, bitter packs a bits<8> into one byte and an instruction struct
into its encoding.
from std.binary import bits
struct Pair {
pub hi: bits<4>,
pub lo: bits<4>,
}
macro pair(hi: int, lo: int) {
@emit Pair(hi as bits<4>, lo as bits<4>)
}
pair 0xA, 0xB
Pair { hi: bits<4> { value: 10 }, lo: bits<4> { value: 11 } }
ab
Emitting from nested calls
A macro’s output includes everything emitted by the macros it calls, in order. So an instruction macro can be built from smaller ones.
A call used as an expression can only emit when its values have somewhere to go. See Where emitted values can go.
Restricting what a macro emits: emits
The emits facet declares which types a macro may emit. It’s optional, but
if it’s there, it’s enforced: emitting anything else is a compile error.
Give several emits facets to allow several types, and use a
wildcard to allow any instance of a generic
type, as in | emits bits<...>. This is how instruction macros declare what
they encode to; -> is only for what a macro returns.
from std.binary import bits
macro byte(value: int)
| emits bits<8>
{
@emit value as bits<8>
}
byte 0x41
bits<8> { value: 65 }
from std.binary import bits
macro byte(value: int)
| emits bits<8>
{
@emit value
}
byte 0x41
`@emit`ed value has type `int`, but this macro's `emits` facet(s) only declare `bits<8>`
A macro with no emits facet may emit anything.
Emitting moves labels
Each emitted value takes up one position in the output, and a label is the position of the next value emitted after it.
@return
@return ends a macro and, optionally, gives a value back to the caller.
Syntax
@return value
@return
Example
macro clamp(x: int, lo: int, hi: int) -> int {
@if x < lo {
@return lo
}
@if x > hi {
@return hi
}
@return x
}
macro show(value: int) {
@emit value
}
show clamp(-5, 0, 10)
show clamp(50, 0, 10)
show clamp(7, 0, 10)
0 10 7
Details
@returnstops the macro at once, even from inside@if,@match,@foror@fold.- A bare
@returnends the macro without a value. - Values the macro emitted before returning stay emitted.
@return f(...), wherefis the macro itself, is a tail call: it doesn’t count toward the recursion limit. See Recursion.@returncan return the value of a call that emits, and its emitted values are kept. See Returning values.
Returning early
macro first_multiple_of(k: int, limit: int) -> int {
@for i in 1..limit {
@if i % k == 0 {
@return i
}
}
@return -1
}
macro show(value: int) {
@emit value
}
show first_multiple_of(7, 100)
show first_multiple_of(700, 100)
7 -1
@assert
@assert stops compilation with an error if a condition is false.
Syntax
@assert condition
@assert condition, "message"
The message is optional, and must be a string literal.
Example
macro shift_amount(n: int) {
@assert n >= 0 && n < 32, "shift amount must be 0 to 31"
@emit n
}
shift_amount 5
5
macro shift_amount(n: int) {
@assert n >= 0 && n < 32, "shift amount must be 0 to 31"
@emit n
}
shift_amount 40
shift amount must be 0 to 31
Details
- The condition is true if it’s not
0. - A passing
@assertemits nothing and has no other effect. - The error points at the
@assertthat failed.
Asserts, hooks and invariants
An @assert checks something at one point in one macro. For checks that
apply more broadly:
- To check the arguments of several macros the same way, put the asserts in
one macro and call it from a
beforehook. - To check every value of a type, wherever it’s made, use an invariant.
@if and @else
@if runs its body only when a condition is true. An optional @else body
runs otherwise.
Syntax
@if condition {
...
}
@if condition {
...
} @else {
...
}
@else goes on the same line as the closing }.
Example
macro abs(x: int) -> int {
@if x < 0 {
@return -x
} @else {
@return x
}
}
macro show(value: int) {
@emit value
}
show abs(-4)
show abs(4)
4 4
More than two branches
There’s no @else @if. Nest another @if inside the @else, or use
@match:
macro sign(x: int) {
@if x < 0 {
@emit -1
} @else {
@if x == 0 {
@emit 0
} @else {
@emit 1
}
}
}
sign -5
sign 0
sign 9
-1 0 1
Choosing an encoding
@if is how an instruction picks between encodings. For example, it might
use a short form when an immediate fits:
macro load_imm(value: int) {
@if value in -2048..2048 {
@emit 1 # one instruction
} @else {
@emit 2 # two instructions
}
}
load_imm 100
load_imm 100_000
1 2
The choice is explicit and visible in the macro. BitterASM never swaps in a “better” encoding on its own.
Outside macros
At the top level, @if includes or skips statements, including
declarations:
const DEBUG = 1
macro show(value: int) {
@emit value
}
@if DEBUG {
show 0xdeb
}
3563
In a struct declaration or construction, it includes or skips fields. See Structs.
@match
@match compares a value against a list of patterns and runs the first arm
that matches.
Syntax
@match value {
pattern => { ... }
pattern => { ... }
_ => { ... }
}
Arms are tried from top to bottom. _ matches anything. If no arm matches,
nothing happens. Commas between arms are optional.
Matching values
A pattern that’s an ordinary expression matches if it equals the value:
macro name_length(n: int) {
@match n {
0 => { @emit 4 } # "zero"
1 => { @emit 3 } # "one"
1 + 1 => { @emit 3 } # "two"
_ => { @emit -1 }
}
}
name_length 0
name_length 2
name_length 7
4 3 -1
Matching enums
Matching is most useful with enums. Name a variant by itself, or qualified by its enum:
enum Color {
Red,
Green,
Blue,
}
macro code(c: Color) {
@match c {
Red => { @emit 1 }
Color.Green => { @emit 2 }
_ => { @emit 3 }
}
}
code Color.Red
code Color.Green
code Color.Blue
1 2 3
For a variant with a payload, Variant(name) binds the payload to name for
that arm. Variant(_) matches any payload without binding it.
enum Shape {
Circle: int,
Square: int,
Empty,
}
macro area(s: Shape) {
@match s {
Shape.Circle(r) => { @emit 3 * r * r }
Square(w) => { @emit w * w }
Empty => { @emit 0 }
}
}
area Shape.Circle(2)
area Shape.Square(5)
area Shape.Empty
12 25 0
Variant(expression), where the expression isn’t a plain name, matches only
when the payload equals it. Qualified forms work the same way:
Shape.Circle(r), and for a generic enum, Option<int>.Some(v).
A pattern that isn’t a variant, such as a constant holding an enum value, is
compared with ==:
enum Color {
Red,
Green,
Blue,
}
const FAVORITE = Color.Blue
macro is_favorite(c: Color) {
@match c {
FAVORITE => { @emit 1 }
_ => { @emit 0 }
}
}
is_favorite Color.Blue
is_favorite Color.Red
1 0
Returning from a match
@return inside an arm returns from the whole macro:
from std.option import Option
macro unwrap_or(o: Option<int>, fallback: int) -> int {
@match o {
Option<int>.Some(v) => { @return v }
None => { @return fallback }
}
}
macro show(value: int) {
@emit value
}
show unwrap_or(Option<int>.Some(42), 0)
show unwrap_or(Option<int>.None, 7)
42 7
@for
@for runs its body once for each element of a source, binding the element
to a name.
Syntax
@for name in source {
...
}
Looping over a range
macro squares(n: int) {
@for i in 0..n {
@emit i * i
}
}
squares 5
0 1 4 9 16
See Ranges for .. and ..=.
Looping over a struct
A struct value works as a source too. @for visits its pub fields in
declaration order, skipping fields marked skip. That’s how you loop over an
array or a string:
from std.array import Array
macro sum<const N: int>(values: Array<int, N>) -> int {
const total = @fold acc = 0 @for v in values {
@next acc + v
}
@return total
}
macro each_char() {
@for c in "hi" {
@emit c
}
}
macro show(value: int) {
@emit value
}
show sum(Array<int, 3> { __el0: 1, __el1: 2, __el2: 3 })
each_char
6 104 105
See Struct fields: pub and skip.
Details
- Each iteration is fresh. Nothing carries over from one iteration to the
next: a
constin the body only exists for that iteration. To carry a value along, such as a running total, use@fold. @returnends the whole macro, not just the loop.- At most 1,000,000 iterations.
- At the top level, the source must be a range written in place, such as
0..32. The loop is unrolled before anything else in the file is evaluated, so it can generate declarations:
@for i in 0..4 {
pub const r`i` = i * 10
}
macro show(value: int) {
@emit value
}
show r3
30
That’s how architecture packages declare their registers. The backticks in
r`i` build each name; see Spliced names.
In struct declarations and constructions
@for can also generate a struct’s fields, or the values of a construction.
See Structs.
@fold and @next
@fold is a @for that carries values from one iteration to the
next. These values are called accumulators. @next ends an iteration
and gives the accumulators their values for the next one.
Syntax
@fold acc = initial @for x in source {
...
@next new_value
}
@fold a = 0, b = 0 @for x in source {
...
@next a = new_a, b = new_b
}
Example: a running total
macro sum_to(n: int) -> int {
@return @fold total = 0 @for i in 0..=n {
@next total + i
}
}
macro show(value: int) {
@emit value
}
show sum_to(10)
55
The fold starts with total = 0. Each iteration sees the current total,
and @next total + i gives the next iteration its new total. The fold’s
value is the final total.
Example: offsets into a table
A fold can emit, too. This one emits each entry’s offset, then the table’s total size:
from std.array import Array
macro offsets<const N: int>(lengths: Array<int, N>) {
const table_size = @fold offset = 0 @for len in lengths {
@emit offset
@next offset + len
}
@emit table_size
}
offsets Array<int, 3> { __el0: 4, __el1: 2, __el2: 5 }
0 4 6 11
Several accumulators
Separate accumulators with commas. @next then names each one it changes.
The fold’s value is a struct with one field per accumulator:
macro stats(n: int) {
const r = @fold count = 0, evens = 0 @for i in 0..n {
@if i % 2 == 0 {
@next count = count + 1, evens = evens + 1
}
@next count = count + 1
}
@emit r.count
@emit r.evens
}
stats 5
5 3
How @next works
@nextends the iteration, likecontinuein other languages. Code after it in the same iteration doesn’t run.- Accumulators
@nextdoesn’t name keep their values, and an iteration that reaches no@nextat all keeps every value. So filtering needs no@else:@if keep { @next total + x }. - Nothing is mutated. Each iteration binds fresh values, the same way
@forbinds its loop variable. @nextbelongs to the innermost@foldin the same macro body. Using it inside a plain@fornested in the fold, or in a macro called from the fold, is an error.- A fold whose body has no
@nextat all gets thefold_without_nextwarning, since its accumulators could never change.
Statement or expression
- As an expression, such as a
const’s value or@return’s value, a fold gives its final accumulators. Its emitted values are kept. - As a statement, its value is ignored. That’s what you want when the body only emits.
No depth limit
A fold runs as many iterations as its source has elements, up to @for’s
limit of 1,000,000. That makes it the tool for long computations that
recursion, limited to 32 nested calls or 4,096 tail
calls, can’t handle.
Everywhere @for works
At the top level the fold is unrolled before anything else, like a
top-level @for. The source must be a range written in place, and the
accumulators are integers. const x = @fold ... names the result, which can
then bound a later top-level @for:
const total = @fold acc = 0 @for i in 0..4 {
@next acc + i
}
macro show(value: int) {
@emit value
}
show total
6
In a struct construction, a fold can compute field values:
from std.array import Array
macro prefix_sums() -> Array<int, 4> {
@return Array<int, 4> {
@fold acc = 0 @for i in 0..4 {
__el`i`: acc + i,
@next acc + i
}
}
}
macro show_all<const N: int>(values: Array<int, N>) {
@for v in values {
@emit v
}
}
show_all prefix_sums()
0 1 3 6
In a struct declaration, a fold can compute field names. There, the accumulators must be integers.
Types
BitterASM has exactly one built-in type, int.
Every other type is declared in BitterASM code, usually in a library:
| Kind | Declared with | Example |
|---|---|---|
| Struct | struct | an instruction format, bits<N>, Array<T, N> |
| Enum | enum | Endian { Little, Big }, Option<T> |
| Type alias | type | type Reg = bits<5> |
Types can be generic: bits<8> and Array<int, 4>
are instances of generic structs.
Why types matter
Types exist only at compile time. The compiler uses them to:
- Check arguments. A macro that takes a
Regrejects anything that isn’t aReg. - Pick overloads.
mov rax, rbxandmov rax, 5call differentmovmacros becauserbxand5have different types. See Overloading. - Enforce rules. An invariant such as “fits in 8 bits” or “is even” is checked every time a value of the type is made.
- Describe output. A struct’s fields tell an evaluator like
bitterhow to lay out bits. See Packing bytes withbitter.
No automatic conversions
A value never changes type on its own. To turn one type into another, convert
it explicitly with as:
from std.binary import bits
macro emit_byte(b: bits<8>) {
@emit b
}
emit_byte 65 as bits<8>
bits<8> { value: 65 }
Equality
== and != work on any two values: structs compare field by field, and
enums compare variant and payload.
In this chapter
- Structs: declaring, constructing and reading them.
- Struct fields:
pubandskip: visibility and iteration. - Enums: a value that is one of several variants.
- Type aliases: new names for types, and new types with rules.
- Invariants: rules every value of a type must follow.
- Conversions:
as,toandfrom.
Structs
A struct groups named fields into one value.
Declaring a struct
struct Name {
field: Type,
pub field: Type,
pub field: Type = default,
}
- Fields are separated by commas or newlines; a trailing comma is fine.
pubmakes a field readable from other files, and visible to@for. See Struct fields.= defaultgives a value used when a construction leaves the field out.pub structlets other files import the struct itself.- Facets such as
invariantgo between the name and the{.
Constructing a struct
There are three ways to build a struct value:
struct Point {
pub x: int,
pub y: int = 5,
}
macro emit_point(p: Point) {
@emit p
}
emit_point Point(1, 2) # positional
emit_point Point(x = 3) # named; y uses its default
emit_point Point { x: 4, y: 6 } # braces
Point { x: 1, y: 2 }
Point { x: 3, y: 5 }
Point { x: 4, y: 6 }
Every field without a default must be given a value:
struct Point {
pub x: int,
pub y: int,
}
macro emit_point(p: Point) {
@emit p
}
emit_point Point(1)
`Point` expects 2 argument(s), but 1 were supplied
A generic struct is constructed with braces:
Pair<int> { a: 1, b: 2 }.
Reading fields
Use .field:
struct Point {
pub x: int,
pub y: int,
}
macro show(value: int) {
@emit value
}
const p = Point(3, 4)
show p.x * p.x + p.y * p.y
25
Values are never modified. To “change” a field, build a new struct.
Structs as machine code
When an evaluator like bitter packs a struct, it concatenates its fields,
first field in the most significant bits. That’s how an instruction format is
described. Here’s RISC-V’s R-type format from std.riscv.impl:
pub struct RType {
funct7: Funct7, # bits<7>
rs2: Reg, # bits<5>
rs1: Reg, # bits<5>
funct3: Funct3, # bits<3>
rd: Reg, # bits<5>
opcode: Opcode, # bits<7>
}
A 32-bit instruction is just a struct whose fields add up to 32 bits. See
Packing bytes with bitter.
Generating fields
A struct’s fields can be generated with @for and
@if, using the struct’s generic parameters. This is how
std.array declares an array of any length:
pub struct Array<T, const N: int>
| invariant N >= 0
{
@for i in 0..N {
pub __el`i`: T,
}
pub skip len: int = N,
}
macro show_all<const N: int>(values: Array<int, N>) {
@for v in values {
@emit v
}
}
show_all Array<int, 3> { __el0: 7, __el1: 8, __el2: 9 }
7 8 9
A construction can generate its field values the same way:
from std.array import Array
macro squares<const N: int>(_len: Array<int, N>) -> Array<int, N> {
@return Array<int, N> {
@for i in 0..N {
__el`i`: i * i,
}
}
}
macro show_all<const N: int>(values: Array<int, N>) {
@for v in values {
@emit v
}
}
show_all squares(Array<int, 4> { __el0: 0, __el1: 0, __el2: 0, __el3: 0 })
0 1 4 9
__el`i` builds the field names __el0, __el1, … See
Spliced names.
Struct fields: pub and skip
Two keywords change how a field can be used: pub and skip.
| Field | Read or named from other files | Visited by @for and in |
|---|---|---|
name: T | No | No |
pub name: T | Yes | Yes |
pub skip name: T | Yes | No |
Every field is part of the value either way: all of them are emitted, and
all of them are packed by bitter.
pub: visibility
A field without pub is private to the file (module) that declares the
struct. Other files can’t read it with .field, or supply it by name when
constructing the struct.
Say shapes.basm declares a struct with one public and one private field:
pub struct Circle {
pub radius: int,
area_cache: int,
}
pub macro circle(r: int) -> Circle {
@return Circle(r, 3 * r * r)
}
Another file can read radius, but not area_cache:
from .shapes import circle
macro show(value: int) {
@emit value
}
show circle(2).radius
2
from .shapes import circle
macro show(value: int) {
@emit value
}
show circle(2).area_cache
field `area_cache` of `Circle` is private to the module that declared it
A few things still work with private fields:
- Positional construction, like
Circle(2, 12), from any file, since it names no fields. - The struct’s own code, meaning its invariants, field defaults and conversions. These always run as part of the declaring file, whoever triggered them.
pub: iteration
@for x in value and x in value
only see pub fields. A private field is treated as internal bookkeeping,
not as one of the struct’s elements.
skip
pub skip marks a field that’s fully public, but isn’t one of the
struct’s elements. @for and in pass over it.
std.array’s Array<T, N> is the example. Its len must be readable by
anyone, but a loop over an array should see only its elements:
pub struct Array<T, const N: int>
| invariant N >= 0
{
@for i in 0..N {
pub __el`i`: T,
}
pub skip len: int = N,
}
macro walk<const N: int>(arr: Array<int, N>) {
@for x in arr {
@emit x
}
@emit arr.len
}
walk Array<int, 2> { __el0: 7, __el1: 8 }
7 8 2
Without skip, the loop would visit len as a third element. String
literals work the same way: their len is pub skip.
skip on a private field is allowed, but does nothing, since private fields
are already left out.
Enums
An enum is a value that is exactly one of a fixed list of variants.
Declaring an enum
enum Name {
Variant,
Variant: PayloadType,
}
A variant can carry one value of a given type, its payload, or nothing.
pub enum lets other files import it.
enum Endian {
Little,
Big,
}
enum Operand {
Register: int,
Immediate: int,
None,
}
macro emit_it(o: Operand) {
@emit o
}
emit_it Operand.Register(3)
emit_it Operand.None
Operand.Register(3)
Operand.None
Making a value
Write the enum’s name, a dot and the variant. Add the payload in parentheses:
| Expression | Value |
|---|---|
Endian.Big | the Big variant |
Operand.Immediate(42) | the Immediate variant, carrying 42 |
Using a value
Compare with ==, or take it apart with @match:
enum Operand {
Register: int,
Immediate: int,
None,
}
macro describe(o: Operand) {
@match o {
Register(n) => { @emit 100 + n }
Immediate(v) => { @emit v }
None => { @emit -1 }
}
}
describe Operand.Register(3)
describe Operand.Immediate(42)
describe Operand.None
103 42 -1
Generic enums
An enum can be generic. std.option declares:
pub enum Option<T> {
Some: T,
None
}
Write the type arguments when making a value: Option<int>.Some(42) and
Option<int>.None. They aren’t inferred:
from std.option import Option
macro emit_option(o: Option<int>) {
@emit o
}
emit_option Option.Some(42)
`Option` expects 1 generic argument(s), but 0 were supplied
Enums as settings
An enum can be a const generic parameter, which
makes it a good fit for options that are fixed at compile time. std.string
uses Endian this way: Utf8String<5, Endian.Little>.
Enums and output
An evaluator decides what an emitted enum means. bitter has no layout for
enums in general, so emitting one to bitter is an error, with one
exception: std.bitter.deferred’s Deferred, which it resolves to a number.
To put an enum in machine code, convert it to a bits<N> first.
Type aliases
type declares a new name for a type.
Syntax
type Name = ExistingType
type Name = ExistingType
| invariant condition
pub type lets other files import it.
Plain aliases: just a new name
Without an invariant, an alias is simply another name for its type. The two are interchangeable:
from std.binary import bits
type Reg = bits<5>
macro use_reg(r: Reg) {
@emit r
}
use_reg Reg(3) # construct through the alias
use_reg 4 as bits<5> # a bits<5> is a Reg
bits<5> { value: 3 }
bits<5> { value: 4 }
Architecture packages use plain aliases to make their types
self-documenting. In std.riscv.impl, Reg, Opcode, Funct3 and Imm12
are all aliases of bits<N>.
Aliases with invariants: a new type
With an invariant, an alias becomes a distinct type
with a rule. A value only becomes one through as, which
checks the rule:
from std.binary import bits
type EvenByte = bits<8>
| invariant v % 2 == 0
macro use_even(b: EvenByte) {
@emit b
}
use_even 6 as EvenByte
bits<8> { value: 6 }
A plain bits<8> isn’t an EvenByte, even if its value is even, because
nothing checked it:
from std.binary import bits
type EvenByte = bits<8>
| invariant v % 2 == 0
macro use_even(b: EvenByte) {
@emit b
}
use_even 6 as bits<8>
type mismatch for `b`: expected `EvenByte`, found `bits<8>`
And as rejects values that break the rule:
from std.binary import bits
type EvenByte = bits<8>
| invariant v % 2 == 0
macro use_even(b: EvenByte) {
@emit b
}
use_even 7 as EvenByte
was violated for `EvenByte`
Naming the value in an invariant
In an alias’s invariant, the value being checked can have any name that
isn’t already declared: v above, x in std.ctypes. The compiler takes the
one free name in the condition to mean the value. Using two different free
names is an error.
Aliases of int
A struct produced by as remembers which alias checked it. A plain integer
has nowhere to record that, so an alias of int works differently: wherever
an int is used as one, as an argument, a struct field or a return value,
the alias’s rule is checked right there. std.unsigned’s uint is an
example:
from std.unsigned import uint
macro count(n: uint) {
@emit n
}
count 5
count 6 as uint
5 6
from std.unsigned import uint
macro count(n: uint) {
@emit n
}
count -1
invariant `(x >= 0)` was violated for `uint`
If a macro has overloads for both int and an alias of int, a plain
integer picks the int one.
Generic aliases
An alias can take generic parameters. Each use substitutes its arguments into the alias’s target and invariants:
from std.binary import bits
type Word<const n: int> = bits<n>
type Aligned<const n: int> = bits<16>
| invariant v % n == 0
macro emit_word(w: Word<12>) {
@emit w
}
macro emit_aligned(a: Aligned<4>) {
@emit a
}
emit_word 5 as bits<12>
emit_aligned 8 as Aligned<4>
bits<12> { value: 5 }
bits<16> { value: 8 }
Word<12> is just another name for bits<12>. Aligned<4> and
Aligned<2> are different types, each with its own rule:
from std.binary import bits
type Aligned<const n: int> = bits<16>
| invariant v % n == 0
macro emit_aligned(a: Aligned<4>) {
@emit a
}
emit_aligned 8 as Aligned<2>
expected `Aligned<4>`, found `Aligned<2>`
Layers
An alias can be built on another type that has rules of its own. Converting
with as checks every layer, from the outside in. std.ctypes declares
uint8_t as a bits<8> whose value is at least zero, so as uint8_t checks
x >= 0, then bits<8>’s own rule that the value fits in eight bits:
from std.ctypes import *
macro emit_byte(b: uint8_t) {
@emit b
}
emit_byte 200 as uint8_t
bits<8> { value: 200 }
from std.ctypes import *
macro emit_byte(b: uint8_t) {
@emit b
}
emit_byte (-1) as uint8_t
invariant `(x >= 0)` was violated for `uint8_t`
Note the parentheses in (-1) as uint8_t. as binds more tightly than -,
so -1 as uint8_t would mean -(1 as uint8_t).
Invariants
An invariant is a rule that every value of a type must follow. It’s checked every time a value of the type is made, so a value that exists is known to be valid.
Syntax
struct Name
| invariant condition
{
fields
}
type Name = Type
| invariant condition
A type can have several invariants, and all of them must hold.
On structs
A struct’s invariant can use its fields and its generic parameters. Refer to
a field by its bare name or as source.field; both work.
struct Range
| invariant lo <= hi
{
pub lo: int,
pub hi: int,
}
macro size(r: Range) -> int {
@return r.hi - r.lo
}
macro show(value: int) {
@emit value
}
show size(Range(2, 10))
8
struct Range
| invariant lo <= hi
{
pub lo: int,
pub hi: int,
}
macro size(r: Range) -> int {
@return r.hi - r.lo
}
macro show(value: int) {
@emit value
}
show size(Range(10, 2))
invariant `(lo <= hi)` was violated for `Range`
It’s checked at every construction, in any form: Range(10, 2),
Range(lo = 10, hi = 2), Range { lo: 10, hi: 2 }, or conversion with as.
bits<N> is an invariant
The standard library’s most important type is a struct with one invariant.
From std.binary:
pub struct bits<const width: int>
| invariant fits_inside_width(width, value)
{
pub skip value: int
}
bits<8> is an integer that’s been checked to fit in 8 bits. An instruction
field typed bits<5> therefore can’t be handed a register number of 40.
Sharing checks
The condition can call macros, so a rule used by several types can live in
one place. That’s what fits_inside_width above is:
pub macro fits_inside_width(width: int, value: int) -> bool {
@return value >= 0 && value < (1 << width)
}
On type aliases
An invariant on a type alias turns it into a new type that only
as can produce. The value being checked can have any
free name. See Type aliases.
from std.binary import bits
type Imm12 = bits<12>
| invariant n % 4 == 0
macro emit_offset(offset: Imm12) {
@emit offset
}
emit_offset 64 as Imm12
bits<12> { value: 64 }
Invariants versus @assert
| Invariant | @assert | |
|---|---|---|
| Belongs to | A type | A macro |
| Checked | Whenever a value of the type is made | When that line runs |
| Good for | Rules about what a value is | Rules about one operation’s inputs |
An invariant may use the struct’s private fields, since it always runs as part of the file that declared the struct.
Conversions: as, to and from
Values never change type on their own. as converts a value explicitly:
value as Type
as binds more tightly than any operator, so put a compound expression in
parentheses: (a + b) as T, (-1) as T.
The same conversion happens when a const has a type: const b: bits<8> = 65
means const b = 65 as bits<8>.
What as does
as tries these steps, in order:
- Same type: nothing to do.
- A conversion the types declare, with a
toorfromfacet. See below. - A type alias: check the alias’s invariants, then convert to the type underneath. See Type aliases.
- A struct with one field: wrap the value in that field, converting it to
the field’s type, and check the struct’s invariants. This is how
65 as bits<8>works:bits<8>has a single field,value.
If nothing applies, it’s an error.
from std.binary import bits
macro emit_byte(b: bits<8>) {
@emit b
}
emit_byte 65 as bits<8>
bits<8> { value: 65 }
Declaring conversions: to and from
A struct or type alias can declare conversions with facets, each naming a macro that performs one:
| to f(source), on the type being converted from: how to turn this type into something else.| from f(source), on the type being converted to: how to make this type from something else.
In the facet, source is the value being converted. as T uses a conversion
when its macro returns T and accepts source’s type.
struct Cents
| to cents_to_int(source)
| from dollars_to_cents(source)
{
pub amount: int,
}
macro cents_to_int(c: Cents) -> int {
@return c.amount
}
macro dollars_to_cents(dollars: int) -> Cents {
@return Cents(dollars * 100)
}
macro show(value: int) {
@emit value
}
show (3 as Cents).amount
show Cents(250) as int
300 250
A type can declare any number of each. If more than one conversion fits, the
as is ambiguous, which is an error. std.decimal’s Decimal and
Fraction use this to convert between each other and int.
Converting to a generic type: target
In a conversion to a generic type, target holds the destination type’s
const parameters, by name. as Fixed<4> sees target.scale as 4:
struct Fixed<const scale: int>
| from int_to_fixed(source, target.scale)
{
pub raw: int,
}
macro int_to_fixed(n: int, scale: int) -> Fixed {
@return Fixed<scale> { raw: n << scale }
}
macro emit_fixed(f: Fixed<4>) {
@emit f
}
emit_fixed 3 as Fixed<4>
Fixed<4> { raw: 48 }
std.string uses target the same way, to convert to
Utf8String<len, endian> with the length and byte order the caller asked
for.
Writing conversion macros
A conversion macro is matched like an ordinary call: as passes it the
facet’s arguments, infers its generic parameters, and checks that it returns
the destination type.
- It can be generic, and its return type can use a wildcard:
struct Fixed<const scale: int>
| to fixed_to_int(source)
{
pub raw: int,
}
macro fixed_to_int<const S: int>(f: Fixed<S>) -> int {
@return f.raw >> S
}
macro show(value: int) {
@emit value
}
show (Fixed<4> { raw: 80 }) as int
5
- A conversion whose macro doesn’t accept the source is simply for other
types, and
asmoves on to the next one. - It must declare its return type, or
ascan’t tell what it converts to:
struct Cents
| from dollars_to_cents(source)
{
pub amount: int,
}
macro dollars_to_cents(dollars: int) {
@return Cents(dollars * 100)
}
macro emit_cents(c: Cents) {
@emit c
}
emit_cents 3 as Cents
conversion macro `dollars_to_cents` must declare its return type
Generics
A generic declaration takes parameters in angle brackets, so one
declaration covers a whole family of types or macros. bits<8> and
bits<32> are two types made from one generic struct, bits<const width: int>.
Two kinds of parameter
| Kind | Written | Stands for | Example |
|---|---|---|---|
| Type parameter | T | a type | Pair<T>, used as Pair<int> |
| Const parameter | const N: int | a value known at compile time | bits<const width: int>, used as bits<8> |
Structs, enums, type aliases and macros can all be generic.
Generic structs
struct Pair<T> {
pub a: T,
pub b: T,
}
macro emit_pair(p: Pair<int>) {
@emit p
}
emit_pair Pair<int> { a: 1, b: 2 }
Pair<int> { a: 1, b: 2 }
A generic struct is constructed with braces, naming its arguments:
Pair<int> { ... }. The Name(...) form is only for non-generic structs.
Generic enums work the same way. See Enums.
Generic macros
A generic macro’s parameters are inferred from its arguments. You never write them at the call:
struct Pair<T> {
pub a: T,
pub b: T,
}
macro first<T>(p: Pair<T>) -> T {
@return p.a
}
macro show(value: int) {
@emit value
}
show first(Pair<int> { a: 5, b: 6 })
5
Inside the macro, a const parameter is an ordinary value, and a type parameter can be used anywhere a type can.
A generic overload loses to a non-generic one that also fits. See Overloading.
In this chapter
- Const parameters: values in types, like
bits<8>andArray<T, N + 1>. - Wildcards:
...: accepting any argument without naming it. - Macros as parameters: passing a macro to a macro.
Const parameters
A const parameter puts a compile-time value into a type. bits<8> and
bits<16> are different types, because their width differs.
Syntax
struct Name<const N: int> { ... }
macro name<const N: int>(x: Type<N>) { ... }
A const parameter’s type can be int or an enum:
from std.binary import Endian
struct Word<const width: int, const order: Endian> {
pub value: int,
}
macro emit_word(w: Word<16, Endian.Big>) {
@emit w
}
emit_word Word<16, Endian.Big> { value: 1 }
Word<16, 1> { value: 1 }
In the output, an enum argument is recorded as its variant’s position in the
enum: Endian.Big is 1, because it’s Endian’s second variant.
Using the value
Inside the declaration, a const parameter is an ordinary value. A struct can use it in its fields, defaults and invariants, and a macro can use it in its body:
from std.binary import bits
macro width_of<const W: int>(_b: bits<W>) -> int {
@return W
}
macro show(value: int) {
@emit value
}
show width_of(3 as bits<12>)
12
Computing types
A type argument can be any expression, including one built from other
parameters. std.array’s appended returns an array one element longer
than its argument:
pub macro appended<T, const N: int>(
arr: Array<T, N>,
value: T
) -> Array<T, N + 1>
from std.binary import Endian
struct Word<const width: int, const order: Endian> {
pub value: int,
}
macro widen<const N: int>(w: Word<N, Endian.Big>) -> Word<N * 2, Endian.Big> {
@return Word<N * 2, Endian.Big> { value: w.value }
}
macro emit_word(w: Word<32, Endian.Big>) {
@emit w
}
emit_word widen(Word<16, Endian.Big> { value: 3 })
Word<32, 1> { value: 3 }
Generating fields from a parameter
A struct’s fields can depend on its const parameters through
@for and @if. That’s how
Array<T, N> has exactly N elements.
Wildcards: ...
In a parameter’s type, ... accepts any argument in that position without
naming it.
from std.array import Array
macro count(arr: Array<int, ...>) -> int {
@return arr.len
}
macro show(value: int) {
@emit value
}
show count(Array<int, 2> { __el0: 1, __el1: 2 })
show count(Array<int, 3> { __el0: 1, __el1: 2, __el2: 3 })
2 3
... versus a named parameter
Both of these accept an array of any length:
macro count(arr: Array<int, ...>) -> int
macro count<const N: int>(arr: Array<int, N>) -> int
Use a named parameter when the macro needs the value as a type argument,
for example to declare its return type as Array<int, N + 1>. Use ... when
it doesn’t. The value is usually still available from the argument itself,
like arr.len above.
In return types
... in a return type says the macro returns some instance of a generic
type, decided by its body. std.bitfield’s field returns a bits<N> whose
width depends on its arguments:
pub macro field(value: int, hi: int, lo: int) -> bits<...> {
@return bits<hi - lo + 1> { value: slice(value, hi, lo) }
}
Mixing
Wildcards and named parameters can be mixed: Array<T, ...> names the
element type and accepts any length. std.array’s get, updated and
reversed all take an Array<T, ...>.
Macros as parameters
A macro can take another macro as an argument, and call it.
macro double(x: int) -> int {
@return x * 2
}
macro apply_twice<F: Fn(int) -> int>(f: F, x: int) -> int {
@return f(f(x))
}
macro show(value: int) {
@emit value
}
show apply_twice(double, 5)
20
Syntax
macro name<F: Fn(ParamType, ...) -> ReturnType>(f: F, ...) { ... }
Fis an ordinary type parameter, with a bound:Fn(int) -> intdescribes the signature a macro must have to be passed asF.- The parameter
f: Freceives the macro. Call it like any macro:f(x). - Pass a macro by its name, with no parentheses:
apply_twice(double, 5). -> ReturnTypecan be left off:Fn(int).
The bound is checked at the call
Passing a macro whose signature doesn’t match is an error at the call site, before the body runs:
macro add(x: int, y: int) -> int {
@return x + y
}
macro apply<F: Fn(int) -> int>(f: F, x: int) -> int {
@return f(x)
}
macro show(value: int) {
@emit value
}
show apply(add, 5)
expected `Fn(int) -> int`, found `Fn(int, int) -> int`
Example: mapping an array
std.array’s mapped applies a macro to every element:
from std.array import Array, mapped
macro double(x: int) -> int {
@return x * 2
}
macro show_all<const N: int>(values: Array<int, N>) {
@for v in values {
@emit v
}
}
show_all mapped(Array<int, 3> { __el0: 1, __el1: 2, __el2: 3 }, double)
2 4 6
Its declaration reads:
pub macro mapped<T, U, F: Fn(T) -> U>(arr: Array<T, ...>, f: F) -> Array<U, ...>
Limits
- Only a non-generic macro can be passed.
Fn(...)is the only kind of bound. There are no traits or interfaces.- A macro is a compile-time value only. It can be passed around and called, but never emitted.
Splicing
Backticks mean “evaluate this now, and put the result here.”
`expression`
There are two uses:
- Building names.
r`i`withi = 3is the namer3. This is the most common use, covered in Spliced names. - Freezing a value into code that runs later. A declaration generated by a macro is evaluated where it ends up, after the macro’s parameters are gone. A splice evaluates part of it while they still exist. See Generating declarations.
macro declare_scaled(n: int) {
pub const SCALED`n` = `n * 100`
}
declare_scaled 7
macro show(value: int) {
@emit value
}
show SCALED7
700
Here the name SCALED`n` becomes SCALED7, and the value `n * 100`
becomes 700, both while declare_scaled runs.
Where it makes no difference
In ordinary expressions, which are evaluated anyway, a splice changes nothing:
`1 + 2` is the same as 1 + 2.
macro show(value: int) {
@emit value
}
show `1 + 2`
3
A literal backtick
To write a backtick that isn’t a splice, for example in a
custom syntax pattern, escape it: \`.
Spliced names
A name can be built from pieces: r`id` is r followed by the value of
id. With id = 3, it’s the name r3.
@for i in 0..4 {
pub const k`i` = i * i
}
macro show_k(n: int) {
@emit k`n`
}
show_k 3
9
The top-level @for declares k0 to k3, and k`n` in show_k reads
the one named by n.
Where names can be spliced
| Position | Example |
|---|---|
| A declaration’s name | pub const x`i` = ..., struct Word`n` { ... } |
| A struct field’s name | pub __el`i`: T, |
After a . | arr.__el`i` |
| An expression | k`n` |
The pieces can be any number of literal parts and splices, in any order:
WORD`bits`_BYTES is WORD16_BYTES when bits is 16. A splice can
hold any expression: arr.__el`arr.len - 1 - i`.
Arrays are built this way
std.array’s Array<T, N> has fields __el0, __el1, … up to N - 1,
generated with spliced names, and its macros reach them the same way:
from std.array import Array
macro reversed_sum<const N: int>(arr: Array<int, N>) {
@for i in 0..N {
@emit arr.__el`N - 1 - i`
}
}
reversed_sum Array<int, 3> { __el0: 1, __el1: 2, __el2: 3 }
3 2 1
The backtick must touch the name
k`n` is one spliced name. k `n`, with a space, is the name k
followed by a separate splice of n.
Reading private constants
A spliced read finds the current file’s non-pub constants as well as pub
ones.
A spliced name can’t carry values between iterations
A const declared inside a @for body exists only for that iteration.
Spliced names can’t be used to pass a value from one iteration to the next.
Use @fold for that.
Modules
Every .basm file is a module. A module sees its own declarations, plus
exactly what it imports, and nothing else.
Module names
A module is named by its path, with . in place of / and without the
extension:
| File | Module |
|---|---|
std/binary.basm | std.binary |
std/riscv/native.basm | std.riscv.native |
helpers.basm, next to the importing file | .helpers |
A leading . makes the path relative to the importing file’s directory.
Without it, the path is looked up in the search path.
A first import
shapes.basm:
pub struct Point {
pub x: int,
pub y: int,
}
pub macro emit_point(p: Point) {
@emit p
}
main.basm, in the same directory:
from .shapes import Point, emit_point
emit_point Point(1, 2)
Point { x: 1, y: 2 }
Every module has its own namespace
Two modules can declare the same name without conflict. A name always means what it meant in the file that wrote it: a macro imported from a library keeps using the library’s helpers, even if the importing file declares something with the same name.
Libraries and architecture packages
The standard library, std, is made of ordinary modules. So is every
architecture: std.riscv.native is a module that imports
std.riscv.impl, and std.x86_64.nasm builds on std.x86_64.intel.
There’s no special kind of module for an instruction set.
In this chapter
- Imports: the forms of
from ... import, and where modules are found. - Visibility and re-exports:
pub,pub from, and what happens when names collide.
Imports
Syntax
from module import * # every `pub` name in the module
from module import a, b, c # only these names
pub from module import ... # import, and re-export
from std.binary import *
from std.string import validate_ascii, string_from_struct
from .helpers import Pair
Importing names
import * brings in every pub declaration of the module, except pub
labels, which must be imported by name (see
Labels).
Listing names imports just those:
pub const WIDTH = 8
pub const HEIGHT = 4
from .consts import WIDTH
macro show(value: int) {
@emit value
}
show WIDTH
8
Importing a name the module doesn’t export, or that isn’t pub, is an
error. An import that’s never used gets the unused_import warning.
Importing a directory
If the module path names a directory, the listed names are modules in it:
from std.riscv import native means from std.riscv.native import *.
from std.riscv import native
add a0, a1, a2
33 85 c5 00
The search path
A relative path starts with dots. One dot is the importing file’s own directory, and each extra dot goes up one more:
| Path | Found in |
|---|---|
.helpers | the importing file’s directory |
.lib.helpers | its lib subdirectory |
..helpers | its parent directory |
...helpers | two directories up |
Here lib/double.basm imports from its parent directory:
pub const BASE = 21
from ..numbers import BASE
pub const DOUBLED = BASE * 2
from .lib.double import DOUBLED
macro show(value: int) {
@emit value
}
show DOUBLED
42
An absolute path (std.binary) is looked up in these directories, in order:
- The current directory.
- Each directory in
BITTERASM_PATH, which is separated likePATH. ~/.bitterasm, where the installer putsstd.
So a project can override a library by putting its own copy earlier in the search path.
Imports aren’t passed on
If a imports b, a file that imports a doesn’t see b’s names. It
must import b itself, unless a re-exports them with pub from. See
Visibility and re-exports.
Imports only affect names. a’s macros still use b however they’re
called, because a name always means what it meant in the file that wrote it.
What an import brings with it
Importing a macro also brings:
- Its overloads. An imported macro’s overloads merge with same-named overloads from other imports and from the importing file. See Overloading.
- Its syntax. A custom syntax assigned to the macro applies in the importing file too.
Visibility and re-exports
pub
A declaration without pub is private to its file. Put pub in front to
let other files import it:
pub const WIDTH = 8
pub macro nop() { ... }
pub struct Point { ... }
pub enum Endian { ... }
pub type Reg = bits<5>
pub start: # a label
Struct fields have their own pub. See Struct fields.
pub const PUBLIC = 1
const PRIVATE = 2
from .lib import PRIVATE
has no `PRIVATE`
Re-exporting with pub from
pub from ... import imports names and exports them again, as if this
file had declared them. That’s how a dialect presents an instruction set
plus its own syntax as one module:
std.riscv.nativedoespub from .impl import *, and adds conventional syntax.std.x86_64.nasmre-exportsstd.x86_64.intel, which re-exportsstd.x86_64.impl.
A program then imports just the dialect.
pub const BASE = 100
pub macro show(value: int) {
@emit value
}
pub from .base import *
pub const EXTRA = 5
from .extended import *
show BASE + EXTRA
105
When names collide
Your own declarations win. A file’s own declaration hides an imported
one with the same name. So adding a new pub name to a library can’t break
a file that already uses that name.
from .base import show
const BASE = 7
show BASE
7
Two imports of one name are ambiguous, but only if you use it. If two
modules both export BASE, and a file imports both with *, using BASE is
an error that names both modules:
pub const BASE = 200
from .base import *
from .other import *
show BASE
`BASE` is imported from more than one module
Importing by name picks one. A name imported by name takes precedence
over one brought in by *:
from .base import *
from .other import BASE
show BASE
200
Macros are the exception. Same-named macros from different modules don’t collide. Their overloads merge into one set, and each call picks the overload that fits. See Overloading.
Programs
A program’s result is the list of values it emits, in order. Everything in this chapter is about arranging that list:
- Labels name positions in it.
- Sections group it into regions, like code and data.
- Linking combines the lists of several files.
- Executables put a header in front, so the result can run.
A complete program
This is examples/x86_64/hello.basm from the repository, a “Hello, world!”
for Linux on x86-64:
from std.x86_64.nasm import *
from std.formats.elf import *
elf64_executable EM_X86_64, _start
const text = "Hello, World!\n"
section .rodata
msg:
db text
section .text
pub _start:
mov eax, 1 # sys_write
mov edi, 1 # stdout
lea rsi, [rel msg] # buffer
mov edx, text.len # length
syscall
mov eax, 60 # sys_exit
xor edi, edi # status 0
syscall
Every piece of it is ordinary BitterASM:
elf64_executableis a macro fromstd.formats.elfthat emits an ELF header. See Executables.section .rodataandsection .textput the string and the code in separate sections.const textnames the string, sodb textcan emit it andtext.lengives its length.msg:and_start:are labels.pubexports_start.lea,mov,syscalland the rest are macros fromstd.x86_64.nasm, anddbis NASM’s data directive.
Build and run it:
bitter build examples/x86_64/hello.basm -o hello
./hello
Top-level statements run in order
The statements at the top level of a file run from top to bottom. Each call appends whatever it emits. Declarations can appear anywhere, since they’re visible throughout the file.
Labels
A label names a position in the program’s output.
Syntax
name:
pub name:
A label goes on a line of its own. Its name can start with ., which is the
convention for a label that’s only used nearby, like a loop. The dot is just
part of the name: .loop isn’t scoped to anything, and every label name must
be unique in its file.
A label’s value is a position
A label’s value is the index of the next value emitted after it: 0 for the
first emitted value, 1 for the second, and so on. It’s an ordinary int:
macro show(value: int) {
@emit value
}
start:
show 10
show 20
middle:
show start
show middle
show end
end:
10 20 0 2 5
A label can be used before it appears in the file, like end above.
From positions to addresses
A position counts emitted values, not bytes. How many bytes each value takes
up is only known once an evaluator packs them. So byte distances are left to
the evaluator: std.bitter.deferred provides span(a, b), a value that
bitter resolves to the number of bytes between positions a and b.
Wrap it in a Positioned<N> to give the result a width. Here a length byte
is written before the data it measures:
from std.binary import bits
from std.bitter.deferred import Positioned, span
macro db(value: int) {
@emit value as bits<8>
}
macro dw(value: int) {
@emit value as bits<16>
}
macro length_byte(start: int, end: int) {
@emit Positioned<8> { value: span(start, end) }
}
length_byte body, done
body:
db 1
dw 2
done:
03 01 00 02
here() is the position of the value being emitted. A relative branch
encodes span(here(), target), and that’s how every jump and branch in
std.riscv and std.x86_64 works. See
Packing bytes with bitter.
Labels across files
pub exports a label. Another file imports it by name; import * never
brings in labels:
pub double:
from .util import double
macro show(value: int) {
@emit value
}
show double
<double>
The importing file can’t know where double is: that depends on how the
files are laid out when they’re linked. So the value is left unresolved, and
bitter build fills it in. See Linking multiple files.
Sections
A section is a named region of the output, such as code or read-only data.
Syntax
section name
Everything emitted after a section statement belongs to that section, until
the next section statement. A name can be reopened any number of times.
The name means nothing to the compiler, or to bitter: .text, .rodata
and code are all just names.
Sections are grouped
When bitter packs a program, it groups each section’s values together.
Sections appear in the order each was first opened, and anything emitted
before the first section statement comes before all of them:
from std.binary import bits
macro db(value: int) {
@emit value as bits<8>
}
db 1
section .data
db 0xaa
section .text
db 2
section .data
db 0xbb
01 aa bb 02
With several input files, bitter build also joins same-named sections
across files. See Linking multiple files.
Sections inside macros
A section statement inside a macro only lasts until the macro returns. The
caller’s section is then restored, so calling a macro can’t move the
caller’s code by accident:
from std.binary import bits
macro db(value: int) {
@emit value as bits<8>
}
macro stash(value: int) {
section .data
db value
}
section .text
db 1
stash 0xaa
db 2
01 02 aa
A macro whose job is to switch the caller’s section, like a shorthand for
section .data, declares the leaks_section facet:
from std.binary import bits
macro db(value: int) {
@emit value as bits<8>
}
macro data()
| leaks_section
{
section .data
}
section .text
db 1
data
db 0xaa
section .text
db 2
01 02 aa
Linking multiple files
A program can be split across files, and bitter build links them into one
image:
bitter build main.basm util.basm -o program
What linking does
- Each input is compiled on its own.
- Same-named sections are joined across all inputs, in
command-line order: every file’s
.text, then every file’s.data, and so on. - Every
publabel that one file imports and another declares is resolved to its final position. - The result is packed into bytes, and written marked as executable.
Example
main.basm calls a routine in util.basm:
# main.basm
from std.riscv.native import *
from .util import double
section .text
pub _start:
addi a0, zero, 21
jal ra, double
# util.basm
from std.riscv.native import *
section .text
pub double:
add a0, a0, a0
jalr zero, ra, 0
bitter build main.basm util.basm -o prog.bin
prog.bin holds four instructions. jal ra, double is encoded as a jump of
4 bytes forward, to where double landed.
Things to know
- The first input goes first. An executable header
must be emitted by the first file on the command line, before any
sectionstatement. - Positions inside a file are kept correct. A branch to a label in the same file still lands on it after sections from other files are merged in around it.
bitter encodedoesn’t link. It packs a single.emfile, and fails if the file refers to a label in another file.
Positions bitter provides
std.bitter.link declares two labels that bitter resolves against the
whole linked image, the way a linker defines symbols like _end:
from std.bitter.link import image_start, image_end
| Label | Position |
|---|---|
image_start | The first value of the image |
image_end | Just past its last value |
span(image_start, image_end) is the image’s size in bytes, and
span(image_start, label) is a label’s offset into the image. That’s what
an executable header needs.
Executables
bitter build writes its output marked as executable, but adds nothing to
it. An executable file format’s header is BitterASM code that the program
writes itself, like any other data.
Adding a header
Call a header macro first thing in the first input file, before any
section statement, so the header starts the image:
from std.formats.elf import *
elf64_executable EM_X86_64, _start
| Format | Module | Header macro |
|---|---|---|
| ELF, 64-bit | std.formats.elf | elf64_executable EM_X86_64, _start |
| ELF, 32-bit | std.formats.elf | elf32_executable EM_RISCV, _start |
| PE32+ (Windows console) | std.formats.pe | pe64_executable IMAGE_FILE_MACHINE_AMD64, _start |
| Mach-O, 64-bit | std.formats.macho | macho64_executable CPU_TYPE_X86_64, CPU_SUBTYPE_X86_64_ALL, _start |
The last argument is the entry point, a pub label that can be
in any input file.
Optional parameters:
- ELF:
load_address,segment_flags(PF_R,PF_W,PF_X),flags. - PE:
image_base. - Mach-O:
vm_address.
Without a header, bitter build writes a flat binary, like nasm -f bin.
What the headers support
Each format maps the whole image as one segment, readable and executable by default. There’s no dynamic linking, no imports and no relocations.
Writing a format of your own
The headers are built from pieces any other format can use:
std.bitter.link’simage_startandimage_end, positionsbitterresolves against the linked image.span(image_start, image_end)is the image’s size in bytes. See Linking.std.bitter.deferred’s arithmetic:add,sub,band,shrand more, which work on positions that aren’t known untilbitterlays the image out.std.bitter.layout’salign n, which emits zero bytes up to the next multiple ofn.std.bitter.layout’spad_image n, which pads the finished image with zeros to a multiple ofn. It takes no space where it’s written, so a header at the start can still pad the end.
Reading std/formats/elf.basm is a good way to see how they fit together.
Evaluators
The compiler, bitterasm, never produces machine code. It runs a program and
records the values it emits, in order, in a .em file. An evaluator
reads that file and decides what the values mean.
prog.basm ──bitterasm compile──▶ prog.em ──evaluator──▶ output
bitter is the evaluator that ships with BitterASM. It packs values into
bytes: binary machine code, or an executable. See
Packing bytes with bitter.
Why split it this way
The language never assumes a program’s output is binary. A bits<8> is a
struct from the standard library, not a language feature. Only bitter
knows that bits<8> means “eight bits.”
So a different evaluator could read the same .em file and produce
something else entirely: a hex dump for a teaching tool, a listing, words for
a 36-bit machine, trits for a ternary one. The language, and every library
that doesn’t depend on bitter, stays the same.
Which types an evaluator knows
In .em, every struct and enum is identified by its module path and name,
such as std.binary.bits. An evaluator gives meaning to the ids it knows,
and treats everything else however its contract says. bitter knows a
handful of ids from std.binary and std.bitter, and packs any other
struct as the concatenation of its fields.
In this chapter
- The
.emformat: the specification an evaluator reads. - Packing bytes with
bitter: howbitterturns values into bytes.
The .em format
bitterasm compile writes a program’s emitted values to a .em file, and
an evaluator such as bitter reads it. This is the contract between them;
anything that reads .em should follow it.
A .em file is one JSON object:
{
"version": 1,
"requires": ["sections", "extern-labels"],
"module": "spec",
"exports": { "start": 0 },
"entries": [
{ "kind": "Struct", "id": "std.binary.bits",
"args": [{ "kind": "Const", "value": "8" }],
"fields": [["value", { "kind": "Int", "value": "7" }]],
"section": ".text" },
{ "kind": "Enum", "id": "spec.Mode", "args": [], "variant": "Slow",
"payload": { "kind": "Int", "value": "3" }, "section": ".text" },
{ "kind": "Deferred", "module": "spec_dep", "symbol": "far", "section": ".text" }
]
}
versionis1. It changes only when the file’s structure changes incompatibly. A reader must refuse any version it doesn’t know, and must refuse the unversioned plain-list files older compilers wrote.requireslists the language features the program actually uses. A reader must refuse a file that requires a feature it doesn’t know, because ignoring one produces wrong output with no error:sections: some entry has asection. Lay entries out grouped by section name, sections in order of first appearance, entries within a section in file order.extern-labels: some value is aDeferred(see below), which only a linker with the other file’s.emcan resolve.
moduleis the compiled file’s module path: its path relative to the deepest search root containing it, with dots (examples.x86_64.hello). A file under no search root is named relative to the working directory, with one leading.per level up plus one (..shared.util).exportsmaps each top-levelpublabel to its position: how many entries precede it.entriesis the emitted values, in emission order. Each is one of these, tagged bykind, plus an optionalsection:Int:valueis a decimal string, since integers are unbounded.Struct:id, genericargs, andfieldsas[name, value]pairs in declaration order.Enum:id, genericargs,variant, and an optionalpayload.Deferred: the value ofpublabelsymbolin the file whosemoduleis given, not known until link time.
A generic argument is {"kind": "Const", "value": "8"} or
{"kind": "Type", "type_kind": ..., ...}, where the type is
{"type_kind": "Builtin", "name": "int"}, or Struct/Enum with an id
and its own args.
Ids. A struct or enum’s id is its declaring module’s path plus its
name: std.binary.bits, spec.Mode. Ids are unique within one program, and
they’re what an evaluator matches on. Which ids an evaluator gives meaning
to is up to that evaluator. bitter understands std.binary.bits,
std.bitter.byte_order.LittleEndian, and std.bitter.deferred’s
Positioned, Deferred, BinOp and Op, and packs any other struct as
the concatenation of its fields.
A later version-1 file may add top-level fields that a reader can safely
ignore. Anything a reader must understand to produce correct output is
either a new requires feature or a new version.
Packing bytes with bitter
bitter turns each emitted value into bits, and each value’s bits into
whole bytes.
bits<N>: a number with a width
std.binary’s bits<N> is the basic unit: an integer that fits in N bits.
It packs to exactly N bits.
from std.binary import bits
macro db(value: int) {
@emit value as bits<8>
}
macro dw(value: int) {
@emit value as bits<16>
}
db 0x41
dw 0x1234
41 12 34
A bare int has no width, so bitter rejects it. Neither does an enum,
which is rejected too.
Structs: fields in order
Any other struct packs as its fields, concatenated in declaration order. The first field goes in the most significant bits:
from std.binary import bits
struct Nibbles {
hi: bits<4>,
lo: bits<4>,
}
macro nibbles(hi: int, lo: int) {
@emit Nibbles(hi as bits<4>, lo as bits<4>)
}
nibbles 0xA, 0xB
ab
That’s the whole instruction-encoding model. An instruction format is a
struct whose fields add up to the instruction’s width, and bitter knows
nothing about opcodes or registers.
Rounding up to bytes
Each top-level value is padded with zero bits at the top to a whole number of bytes:
from std.binary import bits
macro emit12(value: int) {
@emit value as bits<12>
}
emit12 0xabc
0a bc
Byte order
With nothing else said, bitter writes the most significant byte first
(big-endian). That’s the only order that makes sense for a machine that
isn’t byte-addressed at all.
An architecture that wants little-endian output wraps each value in
std.bitter.byte_order’s LittleEndian<T, width>, and bitter reverses its
bytes. width must match the wrapped value’s width, and be a multiple of 8.
from std.binary import bits
from std.bitter.byte_order import LittleEndian
macro dw_le(value: int) {
@emit LittleEndian<bits<16>, 16> { value: value as bits<16> }
}
dw_le 0x1234
34 12
std.riscv and std.x86_64 wrap every instruction this way.
Values not known yet: Positioned<N>
Some values depend on where things end up, such as the distance to a
branch target. std.bitter.deferred builds such values as a Deferred
expression, and Positioned<N> gives one a width:
| Expression | bitter resolves it to |
|---|---|
here() | The position of the value being packed |
span(a, b) | The number of bytes from position a to position b; negative if b comes first |
add, sub, shr, band, … | Arithmetic on the above |
bitter lays the image out first, then resolves each Positioned<N> and
packs it into N bits. See Labels
for an example.
Only use here() directly inside the span of the value being emitted.
Stored and used by a different value, it silently refers to that other
value’s position instead.
Layout: align and pad_image
std.bitter.layout provides two directives that only bitter can resolve,
because they depend on the final layout:
align nemits zero bytes until the next value starts at a multiple ofnbytes from the start of the image.pad_image npads the end of the finished image with zeros to a multiple ofnbytes, wherever it’s written.
from std.binary import bits
from std.bitter.layout import align
macro db(value: int) {
@emit value as bits<8>
}
db 1
align 4
db 2
01 00 00 00 02
Both must be emitted as whole values, never as fields inside another struct.
Sections
Before packing, bitter groups values by section.
With several input files, bitter build also links them. See
Linking multiple files.
Tools
Besides compiling, bitterasm has commands that help while writing code:
| Command | Page |
|---|---|
bitterasm format | Formatting |
bitterasm check | Diagnostics and lints |
bitterasm doc | Generating docs |
bitterasm expand | Generating declarations |
Both format and the lint settings read an optional bitterasm.toml,
found by searching the file’s directory and then its parents.
An editor language server, bitterasm-lsp, can be installed alongside the
compiler. See Installation.
For the full list of commands, see The toolchain.
Formatting
bitterasm format (or fmt) rewrites .basm files in a consistent style:
bitterasm format program.basm # one file
bitterasm fmt std/ # every .basm file below a directory
bitterasm format --check . # change nothing; fail if anything would change
--check is meant for CI: it exits with a non-zero status if any file isn’t
formatted.
What it changes
- Indentation, following
(),[]and{}, plus one level for the lines under a top-level label. - Facets, each moved to its own indented line.
- Trailing whitespace, runs of blank lines, and the final newline.
- Comments longer than
comment_width, which are wrapped. In doc comments, code blocks, tables and headings are left as written. - Long lines, which are wrapped at commas inside
(),[]and{}. A line with no such comma stays long, since a newline would end the statement.
It keeps every comment and never changes what the code means.
Configuration
The formatter reads bitterasm.toml (or .bitterasm.toml), searching the
file’s directory and then its parents, like rustfmt. Pass
--config path/to/bitterasm.toml to choose one. Every setting is optional:
| Setting | Default | Meaning |
|---|---|---|
indent_width | 4 | Spaces per indentation level. |
hard_tabs | false | Indent with tabs instead of spaces. |
indent_facets | true | Indent | facet lines one level under their declaration. |
facets_on_new_line | true | Move facets written on the declaration’s line onto their own lines. |
indent_label_bodies | true | Indent the lines under a top-level label, up to the next label, section or declaration. |
pub_on_declaration | true | Rewrite an old-style | pub line as pub on the declaration. |
return_type_on_declaration | true | Rewrite an old-style | -> T line as -> T on the declaration. |
collapse_short_multiline_generics | true | Join a generic argument list that was split across lines back onto one, when it fits. |
max_blank_lines | 1 | The most consecutive blank lines kept. |
max_width | 100 | The line width code is wrapped at. |
comment_width | 80 | The line width comments are wrapped at. |
newline_style | "Auto" | "Auto", "Unix" or "Windows" line endings. |
indent_width = 2
max_width = 90
The same file holds lint settings, under
[lints].
Diagnostics and lints
Errors
Every error points at the source that caused it, with a file, line and column, the line itself, and a label:
error: shift amount must be 0 to 31
--> prog.basm:2:5
|
2 | @assert n >= 0 && n < 32, "shift amount must be 0 to 31"
| ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ error occurs here
--diagnostic-format picks the output: terminal (the default), plain, or
json for editors and build tools. --color auto|always|never controls
color; auto also respects the NO_COLOR environment variable.
Lints
Warnings are lints. Each has a name, and a level that decides what happens when it fires.
| Lint | Fires when |
|---|---|
unused_import | An imported name is never used. |
unused_parameter | A macro parameter is never used. Start its name with _ to mark it unused on purpose. |
unused_doc_comments | A ## or #! doc comment doesn’t document anything. See Doc comments. |
unreachable_code | A statement comes after a @return or @next that always runs. |
fold_without_next | A @fold body has no @next, so its accumulators never change. |
generated_declarations | Macros generated declarations. See Generating declarations. |
unfulfilled_lint_expectation | An expect facet’s lint didn’t fire. |
missing_docs | A pub item has no ## doc comment, or a file has no #! block. Allowed by default, and not part of all. See Generating docs. |
Two groups name several at once: unused (unused_import,
unused_parameter and unused_doc_comments) and all.
Levels
| Level | Effect |
|---|---|
allow | Silent. |
expect | Silent, but unfulfilled_lint_expectation fires if the lint doesn’t occur. |
warn | A warning. The default for every lint. |
deny | An error. |
forbid | An error, and no declaration can lower it. |
Setting levels
On a declaration, with facets named after the level. Give one lint, or several in parentheses:
macro compatibility(value: int)
| allow unused_parameter
| deny(unreachable_code, generated_declarations)
{
}
expect documents a lint you know about, and tells you once it’s gone:
macro uses_it(x: int)
| expect unused_parameter
{
@emit x
}
uses_it 3
expected `unused_parameter` warning was not produced
(Warnings count as errors in this book’s examples.)
For a project, in bitterasm.toml:
[lints]
unused = "warn"
unreachable_code = "deny"
On the command line, which overrides the project file:
bitterasm compile program.basm -A unused_parameter
bitterasm compile program.basm -D unreachable_code
-A, -W, -D and -F set allow, warn, deny and forbid. They
apply left to right, so a later flag overrides an earlier one:
-D all -A unused denies every lint except the unused group. A forbidden
lint can’t be lowered again, by a later flag or by a facet.
Checking without compiling
bitterasm check runs every check compile does, including lints, but
writes no .em file. It takes the same options.
Generating docs
bitterasm doc turns a module’s doc comments
into reference pages, and checks the examples in them. The
std reference is built by this command,
exactly as your own modules would be.
Syntax
bitterasm doc <paths>... [-o <dir>] # write pages (default: ./doc)
bitterasm doc <paths>... --summary SUMMARY.md # and list them in an mdBook
bitterasm doc <paths>... --test # compile the examples in doc comments
bitterasm doc <paths>... -o <dir> --check # fail if <dir> is out of date
Each path is a .basm file or a directory to search.
Where pages go
The pages follow the modules’ folders. The folder every module shares is
left off, so documenting myarch writes:
doc/
index.md # myarch: what's in it, each with its summary
util.md # myarch.util
x86/
index.md # myarch.x86's contents
native.md # myarch.x86.native
Each page’s title is the module’s own name, with the path to it underneath:
myarch › x86 › native, linking back to each folder. A module with a
folder of the same name next to it (x86.basm and x86/) becomes that
folder’s index.md, followed by the folder’s contents.
Adding the pages to an mdBook
mdBook only shows pages its SUMMARY.md lists. Put two marker lines under
the entry the pages belong to:
- [Reference](reference/index.md)
<!-- bitterasm doc: begin -->
<!-- bitterasm doc: end -->
and pass --summary. bitterasm doc replaces whatever is between the
markers with the pages, nested by folder and indented like the markers:
bitterasm doc myarch -o book/src/reference --summary book/src/SUMMARY.md
What goes on a page
Almost everything on a page comes from the code. Doc comments only add the prose.
- Macros, one section per name, with a row for each overload: how it’s
written, its parameters, and what it returns (
-> T) or emits (| emits T). - Types: structs with their
pubfields, enums with their variants, and type aliases. - Constants and
publabels.
The syntax column shows how a macro is written in that module. Suppose a module declares an instruction and two dialects re-export it:
## Adds two registers.
pub macro add(rd: int, rs1: int, rs2: int) {
@emit rd + rs1 + rs2
}
pub from .impl import *
pub from .impl import *
syntax add(rd, rs1, rs2) = { $rd$ = $rs1$ + $rs2$ }
The page for native shows add rd, rs1, rs2, and the page for c_like
shows rd = rs1 + rs2. Both work:
from .c_like import *
const x = 1
x = 2 + 3
6
A module’s page includes everything it re-exports with pub from.
Re-exported macros are listed in full, since a dialect may spell them
differently, with a column saying which module declares each overload.
Re-exported types and constants read the same everywhere, so they’re
listed by name with a link to their own module’s page.
When a name has several overloads, each row shows the first paragraph of that overload’s doc, and any longer doc appears in full below the table.
Testing examples
--test compiles every example in the doc comments of the given modules,
following the same rules as this book’s examples (see
Writing the docs). A fence
with no language is BitterASM, so an instruction’s encoding can be checked
right where it’s documented:
## Returns from the current function.
##
## ```
## from myarch.native import *
##
## ret
## ```
##
## ```bytes
## c3
## ```
pub macro ret() | emits Byte { ... }
Examples are compiled from the current directory, so they import your
modules the way a program would. Checking bytes needs bitter, found
next to bitterasm or on PATH. A failure names the example’s line:
error: myarch/native.basm:12: expected the bytes c3, but got c2
Keeping pages up to date
--check writes nothing, and fails if any page in the output directory
differs from what bitterasm doc would write, or belongs to a module that
no longer exists. With --summary, it also fails if the list in
SUMMARY.md is out of date. Run it in CI next to --test:
bitterasm doc myarch --test
bitterasm doc myarch -o book/src/reference --summary book/src/SUMMARY.md --check
Writing never deletes anything: a page left over from a module that’s gone is pointed out, for you to delete.
To make sure everything public is documented, turn on the missing_docs
lint, which is allowed by default. It reports each pub item without a
## comment, and a file without a #! block. A macro counts as
documented when any of its overloads in that file is.
bitterasm check myarch/native.basm -W missing_docs
The standard library
std is written entirely in BitterASM. Nothing in it is built into the
compiler: bits, strings and whole instruction sets are ordinary modules you
could have written yourself, and can read in the repository’s std/
directory.
Core
| Module | Provides |
|---|---|
std.binary | bits<N>, two’s-complement signed<N>, the bool type with true and false, Endian, bit_width, byte_width |
std.bitfield | Bit-range helpers for encoders: mask, bit, slice, field, truncate, place |
std.ctypes | C-style integer types: signed int8_t … int64_t and unsigned uint8_t … uint64_t |
std.unsigned | uint, a non-negative int |
std.array | Array<T, N> and get, updated, reversed, popped, appended, first, last, mapped, enumerate, array_from_struct |
std.string | Packed String, AsciiString and Utf8String, with validation and case conversion |
std.option | Option<T> |
std.enumerated | Enumerated<T>, an index with a value |
std.iter | Range and range(start, stop, step) for stepped ranges |
std.decimal | Decimal and Fraction, with conversions between them and int |
std.math | Integer math (pow, gcd, isqrt, popcount, …) and fixed-point math over Decimal (sqrt, ln, sin, atan2, …) |
For bitter
These define the types the bitter evaluator understands. See
Packing bytes with bitter.
| Module | Provides |
|---|---|
std.bitter.deferred | Deferred values, here(), span, arithmetic on them, and Positioned<N> |
std.bitter.byte_order | LittleEndian<T, width> |
std.bitter.layout | align and pad_image |
std.bitter.link | The image_start and image_end labels |
Executable formats
See Executables.
| Module | Provides |
|---|---|
std.formats.elf | elf64_executable, elf32_executable |
std.formats.pe | pe64_executable |
std.formats.macho | macho64_executable |
Architectures
Each architecture has an impl module with its registers, instruction
formats and instructions, plus one or more dialects that give the
instructions a syntax. Import a dialect. See
Custom syntax.
| Architecture | Dialects |
|---|---|
| x86-64 | std.x86_64.intel, std.x86_64.nasm (Intel plus NASM’s [rel label], db and friends), std.x86_64.att |
| RISC-V (RV32I) | std.riscv.native, std.riscv.c_like (a0 = a1 + a2) |
| WebAssembly | std.wasm.module (modules and sections), over std.wasm.impl |
| PDP-10 | std.pdp10.impl (36-bit words, no byte order) |
from std.riscv.c_like import *
a0 = a1 + a2
33 85 c5 00
The same add instruction as in Your first program,
through the C-like dialect.
Reference
The std reference lists every public declaration of
every module, generated from std’s doc comments by bitterasm doc. See
Generating docs to do the same for your own modules.
std
| Name | Summary |
|---|---|
array | Array<T, N>: N values of type T, and ways to build new arrays from old ones. Every value is immutable, so each operation returns a new array. |
binary | Fixed-width binary values: bits<N>, signed<N>, bool, and byte order. |
bitfield | Bit ranges, written the way ISA manuals write them. slice(x, 10, 5) is the manual’s x[10:5]: hi and lo are inclusive bit numbers, so a one-bit field is (n, n). The shift, the mask and the result’s width all come from those two numbers, so they can’t disagree. |
bitter/ | byte_order, deferred, layout, link |
ctypes | C’s fixed-width integer names. The signed ones are signed<N>, stored in two’s complement; the unsigned ones are bits<N> that can’t be negative. |
decimal | Exact non-integer numbers, for compile-time math: Decimal, a value scaled by a power of ten, and Fraction, a numerator over a denominator. Both convert to and from int with as. |
enumerated | Enumerated<T>: a value paired with its position. |
formats/ | elf, macho, pe |
iter | Stepped ranges of integers, for @for. |
math | Integer math, and fixed-point math on Decimal and Fraction, all evaluated at compile time. |
option | Option<T>: a value that may be missing. |
pdp10/ | impl, tops10 |
riscv/ | c_like, impl, native |
string | Strings packed into one integer. A string literal like "hi" is a struct of code points; string_from_struct encodes it as UTF-8 bytes in a single int, with its length in bytes. |
unsigned | uint, an int that can’t be negative. |
wasm/ | impl, leb128, module |
x86_64/ | att, impl, intel, nasm |
array
std › array
Array<T, N>: N values of type T, and ways to build new arrays from
old ones. Every value is immutable, so each operation returns a new array.
from std.array import *
macro show_all<const N: int>(values: Array<int, N>) {
@for v in values {
@emit v
}
}
macro double(x: int) -> int {
@return x * 2
}
const values = Array<int, 3> { __el0: 10, __el1: 20, __el2: 30 }
show_all reversed(values)
show_all mapped(appended(values, 40), double)
30 20 10
20 40 60 80
An array is built with its elements named __el0, __el1, …, or by an
@for inside the construction. @for v in values visits the elements
in order.
Macros
get
The element at index, counting from 0. An index outside the array is
a compile error.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
get(arr, index) | <T>, arr: Array<T, ...>, index: int | returns T |
updated
A copy of arr with the element at index replaced by value.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
updated(arr, index, value) | <T>, arr: Array<T, ...>, index: int, value: T | returns Array<T, ...> |
reversed
A copy of arr in reverse order.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
reversed(arr) | <T>, arr: Array<T, ...> | returns Array<T, ...> |
popped
A copy of arr without its last element. arr can’t be empty.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
popped(arr) | <T, const N: int>, arr: Array<T, N> | returns Array<T, (N - 1)> |
first
The first element. arr can’t be empty.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
first(arr) | <T>, arr: Array<T, ...> | returns T |
last
The last element. arr can’t be empty.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
last(arr) | <T>, arr: Array<T, ...> | returns T |
mapped
A new array holding f of each element, in order. f is any macro
that takes a T and returns a U.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
mapped(arr, f) | <T, U, F: Fn(T) -> U>, arr: Array<T, ...>, f: F | returns Array<U, ...> |
array_from_struct
An Array<int, N> of a struct’s __el0..__el{len - 1} fields, such
as a String’s characters or a Range’s values.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
array_from_struct source | <S>, source: S |
enumerate
Each element paired with its index, as an Enumerated<T>.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
enumerate(arr) | <T>, arr: Array<T, ...> | returns Array<Enumerated<T>, ...> |
appended
A copy of arr with value added at the end.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
appended(arr, value) | <T, const N: int>, arr: Array<T, N>, value: T | returns Array<T, (N + 1)> |
Types
Array
struct Array<T, const N: int>
N values of type T, in order.
| Field | Type | Description |
|---|---|---|
len | int | The number of elements, N. skip, so @for doesn’t visit it. |
Some of its fields are generated by @for or @if.
binary
std › binary
Fixed-width binary values: bits<N>, signed<N>, bool, and byte
order.
bits<N> is how std gives an int a width. Its value must fit in N
bits, which is checked when the value is built, so an encoder can’t
silently emit a field that overflows.
from std.binary import *
macro show<T>(value: T) {
@emit value
}
show bits<8> { value: 65 }
show bit_width(255)
bits<8> { value: 65 }
8
Macros
signed_from_int
x as a signed<width>. It must fit: from -2^(width - 1) up to
2^(width - 1) - 1.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
signed_from_int(x, width) | x: int, width: int | returns signed<...> |
signed_to_int
The int a signed<width> holds, sign-extended from its top bit.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
signed_to_int(x) | <const width: int>, x: signed<width> | returns int |
fits_inside_width
1 if value fits in width bits (0 <= value < 2^width), else 0.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
fits_inside_width(width, value) | width: int, value: int | returns int |
bit_width
The number of bits needed to write n, which must not be negative. Zero
still takes one bit.
from std.binary import bit_width
macro show(value: int) {
@emit value
}
show bit_width(0)
show bit_width(255)
show bit_width(256)
1 8 9
| Syntax | Parameters | Result | Description |
|---|---|---|---|
bit_width(n) | n: int | returns int |
byte_width
The number of whole bytes needed to write n, which must not be
negative: byte_width(255) is 1 and byte_width(256) is 2.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
byte_width(n) | n: int | returns int |
Types
bits
struct bits<const width: int>
An unsigned value width bits wide: 0 <= value < 2^width.
from std.binary import *
const too_big = bits<4> { value: 16 }
invariant `fits_inside_width(width, value)` was violated for `bits`
| Field | Type | Description |
|---|---|---|
value | int | The value itself. skip, so @for over a bits visits nothing. |
signed
struct signed<const width: int>
A signed value width bits wide, stored in two’s complement:
-2^(width - 1) <= value < 2^(width - 1). Convert an int to one with
as, and back the same way.
from std.binary import *
macro db(value: signed<8>) {
@emit value
}
db (-1) as signed<8>
db 100 as signed<8>
ff 64
from std.binary import *
const too_small = (-129) as signed<8>
value doesn't fit in the signed width
| Field | Type | Description |
|---|---|---|
bits | bits<width> | The value’s two’s-complement bits: -1 is all ones. |
bool
type bool = bits<1>
A single bit: true or false.
Endian
enum Endian
Byte order, for code that can lay out values either way.
| Variant | Payload | Description |
|---|---|---|
Little | Least significant byte first, as on x86 and RISC-V. | |
Big | Most significant byte first. |
Constants
| Constant | Type | Value | Description |
|---|---|---|---|
BITS_PER_BYTE | int | 8 | The number of bits in a byte. |
true | bool | 1 | 1 as a bool. |
false | bool | 0 | 0 as a bool. |
bitfield
std › bitfield
Bit ranges, written the way ISA manuals write them. slice(x, 10, 5) is
the manual’s x[10:5]: hi and lo are inclusive bit numbers, so a
one-bit field is (n, n). The shift, the mask and the result’s width all
come from those two numbers, so they can’t disagree.
from std.bitfield import *
macro show(value: int) {
@emit value
}
show slice(0b1101_0110, 7, 4)
show place(3, 7, 6) | place(2, 5, 3) | place(1, 2, 0)
13 209
slice and field also take a Deferred value, such as a branch offset
bitter works out later, and give back a Deferred or Positioned<N>.
Macros
mask
width one-bits: mask(5) is 0b11111.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
mask(width) | width: int | returns int |
bit
Bit n of value, as 0 or 1.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
bit(value, n) | value: int, n: int | returns int |
slice
| Syntax | Parameters | Result | Description |
|---|---|---|---|
slice(value, hi, lo) | value: int, hi: int, lo: int | returns int | Bits hi down to lo of value, shifted down to bit 0: the manuals’ value[hi:lo]. |
slice(value, hi, lo) | value: Deferred, hi: int, lo: int | returns Deferred | Bits hi down to lo of a value bitter works out later. |
field
| Syntax | Parameters | Result | Description |
|---|---|---|---|
field(value, hi, lo) | value: int, hi: int, lo: int | returns bits<...> | slice, as the bits<hi - lo + 1> a format struct’s field holds: imm10_5: field(imm, 10, 5). |
field(value, hi, lo) | value: Deferred, hi: int, lo: int | returns Positioned<...> | slice of a value bitter works out later, as a Positioned<hi - lo + 1>: field(offset, 10, 5) in place of Positioned<6> { value: band(shr(offset, 5), 0b111111) }. |
truncate
The low width bits of value, as a bits<width>. It truncates rather
than rejecting, so a negative immediate becomes its two’s-complement
encoding: truncate(-1, 12) is 0xFFF, where bits<12> { value: -1 }
would fail bits’s range check.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
truncate(value, width) | value: int, width: int | returns bits<...> |
place
value’s low hi - lo + 1 bits, moved up to bits hi down to lo:
slice’s inverse, for building a word out of fields with |.
place(mod, 7, 6) | place(reg, 5, 3) | place(rm, 2, 0) is a ModRM byte.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
place(value, hi, lo) | value: int, hi: int, lo: int | returns int |
bitter
std › bitter
| Name | Summary |
|---|---|
byte_order | Byte order for bitter. Its default is most significant byte first; LittleEndian asks for the reverse. |
deferred | Values that depend on where things end up, such as the distance to a branch target. They’re built as Deferred expressions that bitter resolves once the image is laid out, and given a width by Positioned<N>. |
layout | Padding that depends on the final layout, which only bitter knows: align and pad_image. Emit them as whole values, never as fields inside another struct. |
link | Positions bitter defines when it links a program, the way a linker defines _end. span(image_start, image_end) is the image’s size in bytes, and span(image_start, label) is label’s offset into it, however many files the program spans: what an executable header needs. |
byte_order
Byte order for bitter. Its default is most significant byte first;
LittleEndian asks for the reverse.
from std.binary import bits
from std.bitter.byte_order import LittleEndian
macro dw_le(value: int) {
@emit LittleEndian<bits<16>, 16> { value: value as bits<16> }
}
dw_le 0x1234
34 12
Types
LittleEndian
struct LittleEndian<T, const width: int>
value, packed with its least significant byte first. width is
value’s width in bits, and must be a multiple of 8.
| Field | Type | Description |
|---|---|---|
value | T | The value to reverse, such as a whole instruction. |
deferred
Values that depend on where things end up, such as the distance to a
branch target. They’re built as Deferred expressions that bitter
resolves once the image is laid out, and given a width by
Positioned<N>.
from std.binary import bits
from std.bitter.deferred import *
macro db(value: int) {
@emit value as bits<8>
}
macro offset_to(target: int) {
@emit Positioned<8> { value: span(here(), target) }
}
offset_to end
db 0xAA
db 0xBB
end:
03 aa bb
add, sub, mul, shr and band work on ints, giving an int,
and on Deferreds, giving a Deferred, so an encoder is written the
same way whether its operand is known yet or not.
Macros
here
The position of the value being packed. Use it only inside the span
of the value being emitted: stored and used by a different value, it
refers to that value’s position instead.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
here() | returns Deferred |
add
| Syntax | Parameters | Result | Description |
|---|---|---|---|
add(a, b) | a: int, b: int | returns int | a + b: an int when both are ints, else a Deferred. |
add(a, b) | a: Deferred, b: int | returns Deferred | |
add(a, b) | a: int, b: Deferred | returns Deferred | |
add(a, b) | a: Deferred, b: Deferred | returns Deferred |
sub
| Syntax | Parameters | Result | Description |
|---|---|---|---|
sub(a, b) | a: int, b: int | returns int | a - b: an int when both are ints, else a Deferred. |
sub(a, b) | a: Deferred, b: int | returns Deferred | |
sub(a, b) | a: int, b: Deferred | returns Deferred | |
sub(a, b) | a: Deferred, b: Deferred | returns Deferred |
mul
| Syntax | Parameters | Result | Description |
|---|---|---|---|
mul(a, b) | a: int, b: int | returns int | a * b: an int when both are ints, else a Deferred. |
mul(a, b) | a: Deferred, b: int | returns Deferred | |
mul(a, b) | a: int, b: Deferred | returns Deferred | |
mul(a, b) | a: Deferred, b: Deferred | returns Deferred |
shr
| Syntax | Parameters | Result | Description |
|---|---|---|---|
shr(a, b) | a: int, b: int | returns int | a >> b: an int when a is an int, else a Deferred. |
shr(a, b) | a: Deferred, b: int | returns Deferred |
band
| Syntax | Parameters | Result | Description |
|---|---|---|---|
band(a, b) | a: int, b: int | returns int | a & b: an int when a is an int, else a Deferred. |
band(a, b) | a: Deferred, b: int | returns Deferred |
span
| Syntax | Parameters | Result | Description |
|---|---|---|---|
span(start, end) | start: int, end: int | returns Deferred | The number of bytes from position start to position end, negative if end comes first. Either may be here(). Labels are positions, so span(here(), target) is a relative branch offset. |
span(start, end) | start: Deferred, end: int | returns Deferred | |
span(start, end) | start: int, end: Deferred | returns Deferred |
Types
Op
enum Op
The operation of a BinOp.
| Variant | Payload | Description |
|---|---|---|
Add | left + right. | |
Sub | left - right. | |
Mul | left * right. | |
Shr | left >> right. | |
Band | left & right. | |
Span | Bytes from position left to position right, as in span. |
BinOp
struct BinOp
An operation on two Deferred values.
| Field | Type | Description |
|---|---|---|
op | Op | What to do. |
left | Deferred | The first operand. |
right | Deferred | The second operand. |
Deferred
enum Deferred
A value bitter works out after laying out the image. Build one with
here, span and the arithmetic macros rather than by hand.
| Variant | Payload | Description |
|---|---|---|
Leaf | int | A plain number. |
Pos | int | A position, such as a label, as a number of entries. |
Here | The position of the value being packed. | |
Node | BinOp | An operation on two Deferred values. |
Positioned
struct Positioned<const width: int>
A Deferred value, packed into width bits once bitter resolves it:
what bits<N> is to an int. A negative value is packed in two’s
complement.
| Field | Type | Description |
|---|---|---|
value | Deferred | The value to resolve. |
layout
Padding that depends on the final layout, which only bitter knows:
align and pad_image. Emit them as whole values, never as fields
inside another struct.
from std.binary import bits
from std.bitter.layout import align
macro db(value: int) {
@emit value as bits<8>
}
db 1
align 4
db 2
01 00 00 00 02
Macros
align
Pads with zeros so the next value starts at a multiple of n bytes from
the start of the image.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
align n | n: int |
pad_image
Pads the finished image with zeros to a multiple of n bytes, wherever
it’s written, for formats whose size must be rounded up (PE rounds its
sections to 512 bytes). It takes no space where it appears, so span and
std.bitter.link’s image_end measure the image without the padding.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
pad_image n | n: int |
Types
Align
struct Align<const n: int>
Zero bytes up to the next multiple of n bytes from the start of the
image. align emits one. A span across it counts the padding.
PadImage
struct PadImage<const n: int>
Zero bytes at the end of the image, up to a multiple of n bytes.
pad_image emits one.
link
Positions bitter defines when it links a program, the way a linker
defines _end. span(image_start, image_end) is the image’s size in
bytes, and span(image_start, label) is label’s offset into it, however
many files the program spans: what an executable header needs.
from std.bitter.deferred import *
from std.bitter.link import image_start, image_end
macro size_byte() {
@emit Positioned<8> { value: span(image_start, image_end) }
}
size_byte
size_byte
02 02
Labels
| Label | Description |
|---|---|
image_start | The first byte of the linked image. |
image_end | Just past the last byte of the linked image. |
ctypes
std › ctypes
C’s fixed-width integer names. The signed ones are signed<N>, stored
in two’s complement; the unsigned ones are bits<N> that can’t be
negative.
from std.ctypes import *
macro db(value: int8_t) {
@emit value
}
macro dw(value: uint16_t) {
@emit value
}
db (-2) as int8_t
dw 0x1234 as uint16_t
db (((-2) as int8_t) as int + 5) as int8_t
fe 12 34 03
Types
int8_t
type int8_t = signed<8>
An 8-bit signed value, -128 to 127.
uint8_t
type uint8_t = bits<8>
An 8-bit unsigned value, 0 to 255.
int16_t
type int16_t = signed<16>
A 16-bit signed value.
uint16_t
type uint16_t = bits<16>
A 16-bit unsigned value.
int32_t
type int32_t = signed<32>
A 32-bit signed value.
uint32_t
type uint32_t = bits<32>
A 32-bit unsigned value.
int64_t
type int64_t = signed<64>
A 64-bit signed value.
uint64_t
type uint64_t = bits<64>
A 64-bit unsigned value.
decimal
std › decimal
Exact non-integer numbers, for compile-time math: Decimal, a value
scaled by a power of ten, and Fraction, a numerator over a denominator.
Both convert to and from int with as.
from std.decimal import *
macro show(value: int) {
@emit value
}
const half = Fraction(n = 1, d = 2)
const d = fraction_to_decimal(half, 3)
show d.value
show d.scale
show Decimal(value = 42, scale = 0) as int
500 3 42
Macros
decimal_to_int
x as an int. x.scale must be 0.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
decimal_to_int(x) | x: Decimal | returns int |
decimal_from_int
x as a Decimal with scale 0.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
decimal_from_int(x) | x: int | returns Decimal |
fraction_to_int
x as an int. x.d must be 1.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
fraction_to_int(x) | x: Fraction | returns int |
fraction_from_int
x over 1.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
fraction_from_int(x) | x: int | returns Fraction |
decimal_to_fraction
x as x.value over 10^x.scale, not reduced.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
decimal_to_fraction(x) | x: Decimal | returns Fraction |
fraction_to_decimal
x as a Decimal with precision decimal places, rounded toward
zero.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
fraction_to_decimal(x, precision) | x: Fraction, precision: int | returns Decimal |
Types
Infinity
enum Infinity
A signed infinity.
| Variant | Payload | Description |
|---|---|---|
Positive | Positive infinity. | |
Negative | Negative infinity. |
Decimal
struct Decimal
value / 10^scale: Decimal(value = 314, scale = 2) is 3.14. scale
can’t be negative. as int works when scale is 0.
| Field | Type | Description |
|---|---|---|
value | int | The digits, as an integer. |
scale | int | How many of those digits come after the decimal point. |
Fraction
struct Fraction
n / d, where d isn’t 0. as int works when d is 1, and as Decimal keeps 10 decimal places.
| Field | Type | Description |
|---|---|---|
n | int | The numerator. |
d | int | The denominator. |
enumerated
std › enumerated
Enumerated<T>: a value paired with its position.
Types
Enumerated
struct Enumerated<T>
A value and its index, as enumerate in std.array produces.
| Field | Type | Description |
|---|---|---|
index | int | The position, counting from 0. |
value | T | The value at that position. |
formats
std › formats
| Name | Summary |
|---|---|
elf | ELF static executables, for Linux and the BSDs. |
macho | Mach-O 64 executables, for macOS. |
pe | PE32+ (64-bit Windows) console executables. |
elf
ELF static executables, for Linux and the BSDs.
Invoke one header macro first thing in the program’s entry file, before
any section statement, so its bytes start the image. This one is a
complete Linux x86-64 program that exits with status 0:
from std.formats.elf import *
from std.x86_64.nasm import *
elf64_executable EM_X86_64, _start
_start:
mov eax, 60 # exit
xor edi, edi # with status 0
syscall
The header maps the whole image as one segment at load_address and
starts execution at entry, a label in this file or imported from another
one. Its sizes and entry point come from span over the linked image
(std.bitter.link), so they stay correct however the program’s sections
and files are laid out.
What this doesn’t produce: section headers, dynamic linking, or separate
segments per section. The single segment is readable and executable by
default; pass segment_flags (a combination of PF_R, PF_W, PF_X) to
change that. Every multi-byte field is little-endian, as x86-64 and
RISC-V are.
Macros
elf64_executable
A 64-bit ELF header and its one program header, 120 bytes. machine is
EM_X86_64 or EM_RISCV, entry is the label execution starts at, and
flags is the header’s e_flags, which some architectures give meaning.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
elf64_executable machine, entry, load_address, segment_flags, flags | machine: int, entry: int, load_address: int = 0x400000, segment_flags: int = 5, flags: int = 0 |
elf32_executable
A 32-bit ELF header and its one program header, 84 bytes, e.g. for
RV32. The parameters mean what they do for elf64_executable.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
elf32_executable machine, entry, load_address, segment_flags, flags | machine: int, entry: int, load_address: int = 0x10000, segment_flags: int = 5, flags: int = 0 |
Constants
| Constant | Type | Value | Description |
|---|---|---|---|
EM_X86_64 | 0x3E | machine for x86-64. | |
EM_RISCV | 0xF3 | machine for RISC-V. | |
PF_X | 1 | segment_flags: the segment is executable. | |
PF_W | 2 | segment_flags: the segment is writable. | |
PF_R | 4 | segment_flags: the segment is readable. |
macho
Mach-O 64 executables, for macOS.
Invoke the header macro first thing in the program’s entry file, before
any section statement, so its bytes start the image:
from std.formats.macho import *
from std.x86_64.nasm import *
macho64_executable CPU_TYPE_X86_64, CPU_SUBTYPE_X86_64_ALL, _start
_start:
ret
The whole image is one __TEXT segment at vm_address, readable and
executable, and entry (a label in this file or imported from another
one) is where execution starts. A modern macOS kernel refuses an
executable without a dynamic linker even if it never calls a shared
library, so the header names /usr/lib/dyld; current macOS may still
refuse an unsigned binary depending on Gatekeeper policy.
Not verified on macOS: this reproduces the Rust writer it replaced byte
for byte, and llvm-readobj accepts its output.
Macros
macho64_executable
The Mach-O header and its three load commands (the __TEXT segment, the
dynamic linker and the entry point), 160 bytes. entry is the label
execution starts at, and the segment loads at vm_address.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
macho64_executable cpu_type, cpu_subtype, entry, vm_address | cpu_type: int, cpu_subtype: int, entry: int, vm_address: int = 0x100000000 |
Constants
| Constant | Type | Value | Description |
|---|---|---|---|
CPU_TYPE_X86_64 | 0x01000007 | cpu_type for x86-64. | |
CPU_SUBTYPE_X86_64_ALL | 3 | cpu_subtype for any x86-64 processor. |
pe
PE32+ (64-bit Windows) console executables.
Invoke the header macro first thing in the program’s entry file, before
any section statement, so its bytes start the image:
from std.formats.pe import *
from std.x86_64.nasm import *
pe64_executable IMAGE_FILE_MACHINE_AMD64, _start
_start:
ret
The image is the headers (padded to 512 bytes), then everything after
them as one read+execute .text section at RVA 0x1000, padded to 512
bytes at the end. entry is a label in this file or imported from
another one. No imports, exports or relocations: a program that calls
into Windows DLLs needs more than this provides.
Not verified on Windows: this reproduces the Rust writer it replaced byte
for byte, and llvm-readobj accepts its output.
Macros
pe64_executable
The DOS, PE and optional headers plus the .text section header, 368
bytes, padded to 512. entry is the label execution starts at, and the
image loads at image_base.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
pe64_executable machine, entry, image_base | machine: int, entry: int, image_base: int = 0x140000000 |
Constants
| Constant | Type | Value | Description |
|---|---|---|---|
IMAGE_FILE_MACHINE_AMD64 | 0x8664 | machine for x86-64. |
iter
std › iter
Stepped ranges of integers, for @for.
from std.iter import range
macro countdown() {
@for i in range(10, 0, -3) {
@emit i
}
}
countdown
10 7 4 1
Macros
range
A Range from start to stop, not including stop, every step.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
range start, stop, step | start: int, stop: int, step: int = 1 |
Types
Range
struct Range<const start: int, const stop: int, const step: int>
The integers from start up to (or down to) stop, not including it,
every step. range builds one. step can’t be 0, and has to point from
start towards stop.
Some of its fields are generated by @for or @if.
math
std › math
Integer math, and fixed-point math on Decimal and Fraction, all
evaluated at compile time.
from std.math import *
macro show(value: int) {
@emit value
}
macro demo() {
show gcd(12, 18)
show div_floor(-7, 2)
show sqrt(2, 4).value
show sin(PI(10), 6).value
}
demo
6 -4 14142 0
Functions that give a Decimal take a precision: the number of decimal
places to keep, DEFAULT_PRECISION unless given. Results are truncated
to that many places, so sqrt(2, 4) is 1.4142. Angles are in radians.
Re-exports std.decimal.
Macros
abs
| Syntax | Parameters | Result | Description |
|---|---|---|---|
abs(x) | x: int | returns int | The absolute value of x. |
abs(x) | x: Decimal | returns Decimal | The absolute value of x. |
abs(x) | x: Fraction | returns Fraction | The absolute value of x. |
min
The smaller of a and b.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
min(a, b) | a: int, b: int | returns int |
max
The larger of a and b.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
max(a, b) | a: int, b: int | returns int |
sign
| Syntax | Parameters | Result | Description |
|---|---|---|---|
sign(x) | x: int | returns int | -1, 0 or 1, as x is negative, zero or positive. |
sign(x) | x: Decimal | returns int | -1, 0 or 1, as x is negative, zero or positive. |
sign(x) | x: Fraction | returns int | -1, 0 or 1, as x is negative, zero or positive. |
copysign
abs(x) with the sign of y: abs(x) * sign(y), so 0 when y is 0.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
copysign(x, y) | x: int, y: int | returns int |
clamp
x, moved into lo..=hi if it’s outside. lo can’t be more than hi.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
clamp(x, lo, hi) | x: int, lo: int, hi: int | returns int |
pow
| Syntax | Parameters | Result | Description |
|---|---|---|---|
pow(base, exponent) | base: int, exponent: int | returns int | base to the power exponent, which can’t be negative. |
pow(base, exponent, precision) | base: Decimal, exponent: Decimal, precision: int = DEFAULT_PRECISION | returns Decimal | base to the power exponent. base must be positive. |
pow(base, exponent, precision) | base: Decimal, exponent: int, precision: int = DEFAULT_PRECISION | returns Decimal | base to an integer power. A negative exponent needs a nonzero base. |
gcd
The greatest common divisor of a and b, never negative. gcd(0, 0) is 0.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
gcd(a, b) | a: int, b: int | returns int |
lcm
The least common multiple of a and b, never negative.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
lcm(a, b) | a: int, b: int | returns int |
div_floor
a / b rounded down: div_floor(-7, 2) is -4, where / gives -3.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
div_floor(a, b) | a: int, b: int | returns int |
div_ceil
a / b rounded up: div_ceil(7, 2) is 4.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
div_ceil(a, b) | a: int, b: int | returns int |
rem_euclid
The remainder of a / b that’s never negative: rem_euclid(-7, 3) is 2,
where % gives -1.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
rem_euclid(a, b) | a: int, b: int | returns int |
div_euclid
The quotient that goes with rem_euclid, so
div_euclid(a, b) * b + rem_euclid(a, b) == a.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
div_euclid(a, b) | a: int, b: int | returns int |
pow_mod
base^exponent % modulus, without computing base^exponent in full.
exponent can’t be negative, and modulus must be positive.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
pow_mod(base, exponent, modulus) | base: int, exponent: int, modulus: int | returns int |
isqrt
The square root of x, rounded down. x can’t be negative.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
isqrt(x) | x: int | returns int |
popcount
How many bits of x are 1. x can’t be negative.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
popcount(x) | x: int | returns int |
ctz
How many 0 bits come below x’s lowest 1 bit. ctz(0) is 0.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
ctz(x) | x: int | returns int |
clz
How many 0 bits come above x’s highest 1 bit, in a width-bit value.
x must fit in width bits.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
clz(x, width) | x: int, width: int | returns int |
rotate_left
x rotated left by amount bits within a width-bit value: bits
shifted out at the top come back in at the bottom.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
rotate_left(x, amount, width) | x: int, amount: int, width: int | returns int |
rotate_right
x rotated right by amount bits within a width-bit value: bits
shifted out at the bottom come back in at the top.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
rotate_right(x, amount, width) | x: int, amount: int, width: int | returns int |
is_pow_of_two
1 if x is a power of two, else 0. 0 isn’t one.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
is_pow_of_two(x) | x: int | returns int |
next_pow_of_two
The smallest power of two that’s at least x: next_pow_of_two(5) is 8,
and next_pow_of_two(0) is 1.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
next_pow_of_two(x) | x: int | returns int |
floor
| Syntax | Parameters | Result | Description |
|---|---|---|---|
floor(x) | x: int | returns int | The largest integer not above x. |
floor(x) | x: Decimal | returns int | |
floor(x) | x: Fraction | returns int |
ceil
| Syntax | Parameters | Result | Description |
|---|---|---|---|
ceil(x) | x: int | returns int | The smallest integer not below x. |
ceil(x) | x: Decimal | returns int | |
ceil(x) | x: Fraction | returns int |
trunc
| Syntax | Parameters | Result | Description |
|---|---|---|---|
trunc(x) | x: int | returns int | x with its fractional part dropped, rounding toward zero. |
trunc(x) | x: Decimal | returns int | |
trunc(x) | x: Fraction | returns int |
round
| Syntax | Parameters | Result | Description |
|---|---|---|---|
round(x) | x: int | returns int | The nearest integer to x. Halves round away from zero: 2.5 becomes 3, and -2.5 becomes -3. |
round(x) | x: Decimal | returns int | |
round(x) | x: Fraction | returns int |
fract
| Syntax | Parameters | Result | Description |
|---|---|---|---|
fract(x) | x: Decimal | returns Decimal | x - floor(x), always between 0 and 1: fract(-2.5) is 0.5. |
fract(x) | x: Fraction | returns Fraction |
PI
π to precision decimal places, at most 60.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
PI(precision) | precision: int = DEFAULT_PRECISION | returns Decimal |
TAU
τ = 2π to precision decimal places, at most 60.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
TAU(precision) | precision: int = DEFAULT_PRECISION | returns Decimal |
E
e to precision decimal places, at most 60.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
E(precision) | precision: int = DEFAULT_PRECISION | returns Decimal |
SQRT2
√2 to precision decimal places, at most 60.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
SQRT2(precision) | precision: int = DEFAULT_PRECISION | returns Decimal |
LN2
ln 2 to precision decimal places, at most 60.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
LN2(precision) | precision: int = DEFAULT_PRECISION | returns Decimal |
LN10
ln 10 to precision decimal places, at most 60.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
LN10(precision) | precision: int = DEFAULT_PRECISION | returns Decimal |
root
| Syntax | Parameters | Result | Description |
|---|---|---|---|
root(x, degree, precision) | x: Decimal, degree: int, precision: int = DEFAULT_PRECISION | returns Decimal | The degreeth root of x. degree must be positive, and a negative x needs an odd degree: root(-8, 3) is -2. |
root(x, degree, precision) | x: int, degree: int, precision: int = DEFAULT_PRECISION | returns Decimal | |
root(x, degree, precision) | x: Fraction, degree: int, precision: int = DEFAULT_PRECISION | returns Decimal |
sqrt
| Syntax | Parameters | Result | Description |
|---|---|---|---|
sqrt(x, precision) | x: Decimal, precision: int = DEFAULT_PRECISION | returns Decimal | The square root of x, which can’t be negative. |
sqrt(x, precision) | x: int, precision: int = DEFAULT_PRECISION | returns Decimal |
hypot
| Syntax | Parameters | Result | Description |
|---|---|---|---|
hypot(x, y, precision) | x: Decimal, y: Decimal, precision: int = DEFAULT_PRECISION | returns Decimal | sqrt(x^2 + y^2), the length of the hypotenuse. |
hypot(x, y, precision) | x: int, y: int, precision: int = DEFAULT_PRECISION | returns Decimal |
exp
e to the power x.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
exp(x, precision) | x: Decimal, precision: int = DEFAULT_PRECISION | returns Decimal |
exp2
2 to the power x.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
exp2(x, precision) | x: Decimal, precision: int = DEFAULT_PRECISION | returns Decimal |
expm1
exp(x) - 1.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
expm1(x, precision) | x: Decimal, precision: int = DEFAULT_PRECISION | returns Decimal |
ln
The natural logarithm of x, which must be positive.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
ln(x, precision) | x: Decimal, precision: int = DEFAULT_PRECISION | returns Decimal |
log2
The base-2 logarithm of x, which must be positive.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
log2(x, precision) | x: Decimal, precision: int = DEFAULT_PRECISION | returns Decimal |
log10
The base-10 logarithm of x, which must be positive.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
log10(x, precision) | x: Decimal, precision: int = DEFAULT_PRECISION | returns Decimal |
log1p
ln(1 + x).
| Syntax | Parameters | Result | Description |
|---|---|---|---|
log1p(x, precision) | x: Decimal, precision: int = DEFAULT_PRECISION | returns Decimal |
sin
The sine of x radians.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
sin(x, precision) | x: Decimal, precision: int = DEFAULT_PRECISION | returns Decimal |
cos
The cosine of x radians.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
cos(x, precision) | x: Decimal, precision: int = DEFAULT_PRECISION | returns Decimal |
tan
The tangent of x radians. It’s an error where the cosine is 0.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
tan(x, precision) | x: Decimal, precision: int = DEFAULT_PRECISION | returns Decimal |
atan
The arctangent of x, in radians from -π/2 to π/2.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
atan(x, precision) | x: Decimal, precision: int = DEFAULT_PRECISION | returns Decimal |
atan2
The angle of the point (x, y) from the positive x axis, in radians
from -π to π, using both signs to pick the quadrant. atan2(0, 0) is an
error.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
atan2(y, x, precision) | y: Decimal, x: Decimal, precision: int = DEFAULT_PRECISION | returns Decimal |
asin
The arcsine of x, in radians. x must be between -1 and 1.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
asin(x, precision) | x: Decimal, precision: int = DEFAULT_PRECISION | returns Decimal |
acos
The arccosine of x, in radians. x must be between -1 and 1.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
acos(x, precision) | x: Decimal, precision: int = DEFAULT_PRECISION | returns Decimal |
sinh
The hyperbolic sine of x.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
sinh(x, precision) | x: Decimal, precision: int = DEFAULT_PRECISION | returns Decimal |
cosh
The hyperbolic cosine of x.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
cosh(x, precision) | x: Decimal, precision: int = DEFAULT_PRECISION | returns Decimal |
tanh
The hyperbolic tangent of x.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
tanh(x, precision) | x: Decimal, precision: int = DEFAULT_PRECISION | returns Decimal |
asinh
The inverse hyperbolic sine of x.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
asinh(x, precision) | x: Decimal, precision: int = DEFAULT_PRECISION | returns Decimal |
acosh
The inverse hyperbolic cosine of x, which must be at least 1.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
acosh(x, precision) | x: Decimal, precision: int = DEFAULT_PRECISION | returns Decimal |
atanh
The inverse hyperbolic tangent of x, which must be between -1 and 1,
exclusive.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
atanh(x, precision) | x: Decimal, precision: int = DEFAULT_PRECISION | returns Decimal |
simplify
x in lowest terms, with a positive denominator: 6/-4 becomes -3/2.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
simplify(x) | x: Fraction | returns Fraction |
Constants
| Constant | Type | Value | Description |
|---|---|---|---|
DEFAULT_PRECISION | int | 16 | The number of decimal places a Decimal result keeps when no precision is given. |
GUARD_DIGITS | int | 4 | Extra decimal places used inside a calculation, and dropped from its result. |
CONSTANT_DIGITS | int | 60 | The most decimal places PI, E and the other constants can give. |
Re-exported
From std.decimal: Decimal, Fraction.
option
std › option
Option<T>: a value that may be missing.
Types
Option
enum Option<T>
Either Some value of type T, or None.
| Variant | Payload | Description |
|---|---|---|
Some | T | A value. |
None | No value. |
pdp10
std › pdp10
| Name | Summary |
|---|---|
impl | The PDP-10: 36-bit words, 16 accumulators, and a representative subset of its instructions, all in the one word format. |
tops10 | TOPS-10 monitor calls: the operating system’s side of a PDP-10 program, kept apart from std.pdp10.impl, which is only the machine. |
impl
The PDP-10: 36-bit words, 16 accumulators, and a representative subset of its instructions, all in the one word format.
from std.pdp10.impl import *
movei ac1, 0, Index(0), 42
04 08 80 00 2a
Each instruction takes its operands as the manual’s fields: accumulator
ac, indirect bit i, index register x, and address y. The machine
is word-addressed, so a word has no byte order; bitter packs its 36
bits into 5 bytes with 4 zero bits above them.
Opcodes are from the DEC PDP-10 System Reference Manual (DEC-10-HGAA-D), Appendix A.
Macros
instr
An instruction with any opcode, for one this module doesn’t name, such as an operating system’s monitor call.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
instr opcode, ac, i, x, y | opcode: int, ac: int, i: int, x: int, y: int | emits Instr |
halt
Stops the processor. It’s JRST 4,.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
halt | emits Instr |
jrst
Jumps to the effective address.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jrst i, x, y | i: int, x: Index, y: int | emits Instr |
jumpa
Jumps to the effective address. JUMPA is JUMP with the “always”
condition.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jumpa i, x, y | i: int, x: Index, y: int | emits Instr |
movei
Loads the effective address itself, not what’s there, into ac: an
immediate load.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
movei ac, i, x, y | ac: AC, i: int, x: Index, y: int | emits Instr |
move
Loads the word at the effective address into ac.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
move ac, i, x, y | ac: AC, i: int, x: Index, y: int | emits Instr |
movem
Stores ac at the effective address.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
movem ac, i, x, y | ac: AC, i: int, x: Index, y: int | emits Instr |
add
Adds the word at the effective address into ac. Addresses 0 to 15 are
the accumulators, so add(ac1, 0, Index(0), 2) adds ac2 into ac1.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
add ac, i, x, y | ac: AC, i: int, x: Index, y: int | emits Instr |
addi
Adds the effective address itself into ac.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
addi ac, i, x, y | ac: AC, i: int, x: Index, y: int | emits Instr |
sub
Subtracts the word at the effective address from ac.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
sub ac, i, x, y | ac: AC, i: int, x: Index, y: int | emits Instr |
and
ANDs the word at the effective address into ac.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
and ac, i, x, y | ac: AC, i: int, x: Index, y: int | emits Instr |
cain
Compares ac with the effective address itself, and skips the next
instruction if they’re not equal.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
cain ac, i, x, y | ac: AC, i: int, x: Index, y: int | emits Instr |
exch
Swaps ac with the word at the effective address.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
exch ac, i, x, y | ac: AC, i: int, x: Index, y: int | emits Instr |
asciz
A string packed the way MACRO-10’s ASCIZ packs it: five 7-bit
characters to a word, first character in the high bits, low bit unused,
and at least one NUL at the end. Each word is its own emitted value, so a
label after it still counts words.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
asciz source | <S>, source: S |
Types
AC
type AC = bits<4>
An accumulator number, 0 to 15.
Opcode
type Opcode = bits<9>
The 9-bit opcode.
Indirect
type Indirect = bits<1>
The indirect bit: 1 means y holds the address of the operand’s address.
Index
type Index = bits<4>
An index register number. Any accumulator but 0 can be one; 0 means none.
Addr
type Addr = bits<18>
An 18-bit address.
Instr
struct Instr
The PDP-10 instruction word: opcode (9 bits), ac (4), i (1), x (4)
and y (18), most significant first.
Word
type Word = bits<36>
A 36-bit data word.
Constants
| Constant | Type | Value | Description |
|---|---|---|---|
ac0 | AC(0) | An accumulator. None is wired to zero, and there are no ABI names: each PDP-10 operating system had its own conventions. | |
ac1 | AC(1) | An accumulator. None is wired to zero, and there are no ABI names: each PDP-10 operating system had its own conventions. | |
ac2 | AC(2) | An accumulator. None is wired to zero, and there are no ABI names: each PDP-10 operating system had its own conventions. | |
ac3 | AC(3) | An accumulator. None is wired to zero, and there are no ABI names: each PDP-10 operating system had its own conventions. | |
ac4 | AC(4) | An accumulator. None is wired to zero, and there are no ABI names: each PDP-10 operating system had its own conventions. | |
ac5 | AC(5) | An accumulator. None is wired to zero, and there are no ABI names: each PDP-10 operating system had its own conventions. | |
ac6 | AC(6) | An accumulator. None is wired to zero, and there are no ABI names: each PDP-10 operating system had its own conventions. | |
ac7 | AC(7) | An accumulator. None is wired to zero, and there are no ABI names: each PDP-10 operating system had its own conventions. | |
ac8 | AC(8) | An accumulator. None is wired to zero, and there are no ABI names: each PDP-10 operating system had its own conventions. | |
ac9 | AC(9) | An accumulator. None is wired to zero, and there are no ABI names: each PDP-10 operating system had its own conventions. | |
ac10 | AC(10) | An accumulator. None is wired to zero, and there are no ABI names: each PDP-10 operating system had its own conventions. | |
ac11 | AC(11) | An accumulator. None is wired to zero, and there are no ABI names: each PDP-10 operating system had its own conventions. | |
ac12 | AC(12) | An accumulator. None is wired to zero, and there are no ABI names: each PDP-10 operating system had its own conventions. | |
ac13 | AC(13) | An accumulator. None is wired to zero, and there are no ABI names: each PDP-10 operating system had its own conventions. | |
ac14 | AC(14) | An accumulator. None is wired to zero, and there are no ABI names: each PDP-10 operating system had its own conventions. | |
ac15 | AC(15) | An accumulator. None is wired to zero, and there are no ABI names: each PDP-10 operating system had its own conventions. |
tops10
TOPS-10 monitor calls: the operating system’s side of a PDP-10 program,
kept apart from std.pdp10.impl, which is only the machine.
from std.pdp10.impl import *
from std.pdp10.tops10 import *
start:
outstr JOBDA + msg
exit
msg:
asciz "Hi\r\n"
01 49 80 00 62
01 38 00 00 0a
09 1a 46 8a 00
TOPS-10 loads a program at JOBDA. Every instruction and data word is
one emitted value, so a label’s position counts words, and JOBDA plus
the label is its address.
Opcodes are from the DECsystem-10 Monitor Calls manual (AA-0974G-TB).
Macros
outstr
Types the ASCIZ string at address on the terminal: TTCALL 3,.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
outstr address | address: int | emits Instr |
outchr
Types the character in the low 7 bits of the word at address on the
terminal: TTCALL 1,.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
outchr address | address: int | emits Instr |
exit
Ends the program and returns to the monitor: CALLI 12.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
exit | emits Instr |
Constants
| Constant | Type | Value | Description |
|---|---|---|---|
JOBDA | 0o140 | .JBDA, the first address after the job data area: where TOPS-10 loads a program’s code. |
riscv
std › riscv
| Name | Summary |
|---|---|
c_like | RISC-V RV32I written as expressions instead of mnemonics: a0 = a1 + a2, a1 = mem[sp + 8], if (a0 != zero) goto loop. |
impl | RISC-V RV32I: its registers, instruction formats, and every base instruction, each emitted as a little-endian 32-bit word. |
native | RISC-V RV32I in its standard assembly syntax: add a0, a1, a2, lw a0, 8(sp), beq a0, zero, done. |
c_like
RISC-V RV32I written as expressions instead of mnemonics:
a0 = a1 + a2, a1 = mem[sp + 8], if (a0 != zero) goto loop.
from std.riscv.c_like import *
loop:
a0 <- a0 + -1
if (a0 != zero) goto loop
a1 = mem[sp + 8]
13 05 f5 ff
e3 1e 05 fe
83 25 81 00
- A register-register operation uses
=, and one with an immediate uses<-:a0 = a0 + a1isadd, anda0 <- a0 + 1isaddi. >>>is a logical shift right, and>>an arithmetic one.- Loads and stores smaller than a word name their width before
mem:i8,u8,i16oru16. - Unsigned comparisons (
sltu,sltiu,bltu,bgeu),lui,auipc,ecall,ebreak,li,laandasciihave no C-like spelling, and keep their plainname a, b, csyntax.
Import this instead of std.riscv.native, never alongside it: both
assign syntax to the same instructions.
Re-exports std.riscv.impl.
Macros
add
rd = rs1 + rs2.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rd = rs1 + rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> | std.riscv.impl |
sub
rd = rs1 - rs2.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rd = rs1 - rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> | std.riscv.impl |
sll
rd = rs1 << rs2, shifting by the low 5 bits of rs2.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rd = rs1 << rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> | std.riscv.impl |
slt
rd = 1 if rs1 < rs2 as signed numbers, else 0.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rd = rs1 < rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> | std.riscv.impl |
sltu
rd = 1 if rs1 < rs2 as unsigned numbers, else 0.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
sltu rd, rs1, rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> | std.riscv.impl |
xor
rd = rs1 ^ rs2.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rd = rs1 ^ rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> | std.riscv.impl |
srl
rd = rs1 >> rs2, shifting in zeros, by the low 5 bits of rs2.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rd = rs1 >>> rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> | std.riscv.impl |
sra
rd = rs1 >> rs2, shifting in copies of the sign bit, by the low 5 bits of
rs2.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rd = rs1 >> rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> | std.riscv.impl |
or
rd = rs1 | rs2.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rd = rs1 | rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> | std.riscv.impl |
and
rd = rs1 & rs2.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rd = rs1 & rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> | std.riscv.impl |
addi
rd = rs1 + imm, with a 12-bit signed imm.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rd <- rs1 + imm | rd: Reg, rs1: Reg, imm: int | emits LittleEndian<IType, 32> | std.riscv.impl |
slti
rd = 1 if rs1 < imm as signed numbers, else 0.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rd <- rs1 < imm | rd: Reg, rs1: Reg, imm: int | emits LittleEndian<IType, 32> | std.riscv.impl |
sltiu
rd = 1 if rs1 < imm as unsigned numbers, else 0. imm is
sign-extended first, so sltiu rd, rs1, 1 tests for zero.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
sltiu rd, rs1, imm | rd: Reg, rs1: Reg, imm: int | emits LittleEndian<IType, 32> | std.riscv.impl |
xori
rd = rs1 ^ imm. xori rd, rs1, -1 is bitwise NOT.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rd <- rs1 ^ imm | rd: Reg, rs1: Reg, imm: int | emits LittleEndian<IType, 32> | std.riscv.impl |
ori
rd = rs1 | imm.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rd <- rs1 | imm | rd: Reg, rs1: Reg, imm: int | emits LittleEndian<IType, 32> | std.riscv.impl |
andi
rd = rs1 & imm.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rd <- rs1 & imm | rd: Reg, rs1: Reg, imm: int | emits LittleEndian<IType, 32> | std.riscv.impl |
slli
rd = rs1 << shamt, for shamt from 0 to 31.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rd <- rs1 << shamt | rd: Reg, rs1: Reg, shamt: int | emits LittleEndian<IType, 32> | std.riscv.impl |
srli
rd = rs1 >> shamt, shifting in zeros, for shamt from 0 to 31.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rd <- rs1 >>> shamt | rd: Reg, rs1: Reg, shamt: int | emits LittleEndian<IType, 32> | std.riscv.impl |
srai
rd = rs1 >> shamt, shifting in copies of the sign bit, for shamt from 0
to 31.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rd <- rs1 >> shamt | rd: Reg, rs1: Reg, shamt: int | emits LittleEndian<IType, 32> | std.riscv.impl |
jalr
Jumps to rs1 + offset (with bit 0 cleared) and puts the address of the
next instruction in rd. jalr zero, ra, 0 returns from a function.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rd = call[rs1 + offset] | rd: Reg, rs1: Reg, offset: int | emits LittleEndian<IType, 32> | std.riscv.impl |
ecall
Calls the execution environment, e.g. a Linux system call.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
ecall | emits LittleEndian<IType, 32> | std.riscv.impl |
ebreak
Stops in a debugger.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
ebreak | emits LittleEndian<IType, 32> | std.riscv.impl |
lb
Loads the byte at rs1 + offset into rd, sign-extended.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rd = i8 mem[rs1 + offset] | rd: Reg, offset: int, rs1: Reg | emits LittleEndian<IType, 32> | std.riscv.impl |
lh
Loads the 16-bit halfword at rs1 + offset into rd, sign-extended.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rd = i16 mem[rs1 + offset] | rd: Reg, offset: int, rs1: Reg | emits LittleEndian<IType, 32> | std.riscv.impl |
lw
Loads the 32-bit word at rs1 + offset into rd.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rd = mem[rs1 + offset] | rd: Reg, offset: int, rs1: Reg | emits LittleEndian<IType, 32> | std.riscv.impl |
lbu
Loads the byte at rs1 + offset into rd, zero-extended.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rd = u8 mem[rs1 + offset] | rd: Reg, offset: int, rs1: Reg | emits LittleEndian<IType, 32> | std.riscv.impl |
lhu
Loads the 16-bit halfword at rs1 + offset into rd, zero-extended.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rd = u16 mem[rs1 + offset] | rd: Reg, offset: int, rs1: Reg | emits LittleEndian<IType, 32> | std.riscv.impl |
sb
Stores the low byte of rs2 at rs1 + offset.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
i8 mem[rs1 + offset] = rs2 | rs2: Reg, offset: int, rs1: Reg | emits LittleEndian<SType, 32> | std.riscv.impl |
sh
Stores the low 16 bits of rs2 at rs1 + offset.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
i16 mem[rs1 + offset] = rs2 | rs2: Reg, offset: int, rs1: Reg | emits LittleEndian<SType, 32> | std.riscv.impl |
sw
Stores rs2 at rs1 + offset.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
mem[rs1 + offset] = rs2 | rs2: Reg, offset: int, rs1: Reg | emits LittleEndian<SType, 32> | std.riscv.impl |
beq
Branches to the label target if rs1 == rs2.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
if (rs1 == rs2) goto target | rs1: Reg, rs2: Reg, target: int | emits LittleEndian<BType, 32> | std.riscv.impl |
bne
Branches to the label target if rs1 != rs2.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
if (rs1 != rs2) goto target | rs1: Reg, rs2: Reg, target: int | emits LittleEndian<BType, 32> | std.riscv.impl |
blt
Branches to the label target if rs1 < rs2 as signed numbers.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
if (rs1 < rs2) goto target | rs1: Reg, rs2: Reg, target: int | emits LittleEndian<BType, 32> | std.riscv.impl |
bge
Branches to the label target if rs1 >= rs2 as signed numbers.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
if (rs1 >= rs2) goto target | rs1: Reg, rs2: Reg, target: int | emits LittleEndian<BType, 32> | std.riscv.impl |
bltu
Branches to the label target if rs1 < rs2 as unsigned numbers.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
bltu rs1, rs2, target | rs1: Reg, rs2: Reg, target: int | emits LittleEndian<BType, 32> | std.riscv.impl |
bgeu
Branches to the label target if rs1 >= rs2 as unsigned numbers.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
bgeu rs1, rs2, target | rs1: Reg, rs2: Reg, target: int | emits LittleEndian<BType, 32> | std.riscv.impl |
lui
rd = imm << 12: loads a 20-bit upper immediate.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
lui rd, imm | rd: Reg, imm: int | emits LittleEndian<UType, 32> | std.riscv.impl |
auipc
rd = pc + (imm << 12): an address relative to this instruction.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
auipc rd, imm | rd: Reg, imm: int | emits LittleEndian<UType, 32> | std.riscv.impl |
jal
Jumps to the label target and puts the address of the next instruction
in rd. jal ra, f calls f; jal zero, l just jumps.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rd = call(target) | rd: Reg, target: int | emits LittleEndian<JType, 32> | std.riscv.impl |
li
Loads the constant imm into rd: addi rd, zero, imm when it fits in
12 signed bits, else lui for the upper 20 bits then, unless they’re
zero, addi for the rest.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
li rd, imm | rd: Reg, imm: int | std.riscv.impl |
la
Loads the address of the label symbol into rd: auipc for the upper
20 bits of the distance, then addi for the rest.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
la rd, symbol | rd: Reg, symbol: int | std.riscv.impl |
ascii
A string literal’s UTF-8 bytes, with no terminator: GNU’s .ascii.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
ascii source | <S>, source: S | std.riscv.impl |
Re-exported
From std.riscv.impl: Reg, x0, x1, x2, x3, x4, x5, x6, x7, x8, x9, x10, x11, x12, x13, x14, x15, x16, x17, x18, x19, x20, x21, x22, x23, x24, x25, x26, x27, x28, x29, x30, x31, zero, ra, sp, gp, tp, t0, t1, t2, s0, fp, s1, a0, a1, a2, a3, a4, a5, a6, a7, s2, s3, s4, s5, s6, s7, s8, s9, s10, s11, t3, t4, t5, t6, Opcode, Funct3, Funct7, Bit1, Imm4, Imm5, Imm6, Imm7, Imm8, Imm10, Imm12, Imm20, RType, IType, SType, BType, UType, JType, Byte, Bytes.
impl
RISC-V RV32I: its registers, instruction formats, and every base instruction, each emitted as a little-endian 32-bit word.
Import a dialect rather than this module: std.riscv.native for the
standard assembly syntax, or std.riscv.c_like. Here every instruction
takes its operands in order, lw rd, offset, rs1 rather than
lw rd, offset(rs1).
from std.riscv.impl import *
loop:
addi a0, a0, -1
bne a0, zero, loop
13 05 f5 ff
e3 1e 05 fe
Immediates are plain ints, truncated to their field, so a negative one
becomes its two’s-complement encoding. Branch and jump targets are labels;
bitter works out the offset once the program is laid out.
Macros
add
rd = rs1 + rs2.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
add rd, rs1, rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> |
sub
rd = rs1 - rs2.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
sub rd, rs1, rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> |
sll
rd = rs1 << rs2, shifting by the low 5 bits of rs2.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
sll rd, rs1, rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> |
slt
rd = 1 if rs1 < rs2 as signed numbers, else 0.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
slt rd, rs1, rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> |
sltu
rd = 1 if rs1 < rs2 as unsigned numbers, else 0.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
sltu rd, rs1, rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> |
xor
rd = rs1 ^ rs2.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
xor rd, rs1, rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> |
srl
rd = rs1 >> rs2, shifting in zeros, by the low 5 bits of rs2.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
srl rd, rs1, rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> |
sra
rd = rs1 >> rs2, shifting in copies of the sign bit, by the low 5 bits of
rs2.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
sra rd, rs1, rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> |
or
rd = rs1 | rs2.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
or rd, rs1, rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> |
and
rd = rs1 & rs2.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
and rd, rs1, rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> |
addi
rd = rs1 + imm, with a 12-bit signed imm.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
addi rd, rs1, imm | rd: Reg, rs1: Reg, imm: int | emits LittleEndian<IType, 32> |
slti
rd = 1 if rs1 < imm as signed numbers, else 0.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
slti rd, rs1, imm | rd: Reg, rs1: Reg, imm: int | emits LittleEndian<IType, 32> |
sltiu
rd = 1 if rs1 < imm as unsigned numbers, else 0. imm is
sign-extended first, so sltiu rd, rs1, 1 tests for zero.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
sltiu rd, rs1, imm | rd: Reg, rs1: Reg, imm: int | emits LittleEndian<IType, 32> |
xori
rd = rs1 ^ imm. xori rd, rs1, -1 is bitwise NOT.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
xori rd, rs1, imm | rd: Reg, rs1: Reg, imm: int | emits LittleEndian<IType, 32> |
ori
rd = rs1 | imm.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
ori rd, rs1, imm | rd: Reg, rs1: Reg, imm: int | emits LittleEndian<IType, 32> |
andi
rd = rs1 & imm.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
andi rd, rs1, imm | rd: Reg, rs1: Reg, imm: int | emits LittleEndian<IType, 32> |
slli
rd = rs1 << shamt, for shamt from 0 to 31.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
slli rd, rs1, shamt | rd: Reg, rs1: Reg, shamt: int | emits LittleEndian<IType, 32> |
srli
rd = rs1 >> shamt, shifting in zeros, for shamt from 0 to 31.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
srli rd, rs1, shamt | rd: Reg, rs1: Reg, shamt: int | emits LittleEndian<IType, 32> |
srai
rd = rs1 >> shamt, shifting in copies of the sign bit, for shamt from 0
to 31.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
srai rd, rs1, shamt | rd: Reg, rs1: Reg, shamt: int | emits LittleEndian<IType, 32> |
jalr
Jumps to rs1 + offset (with bit 0 cleared) and puts the address of the
next instruction in rd. jalr zero, ra, 0 returns from a function.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jalr rd, rs1, offset | rd: Reg, rs1: Reg, offset: int | emits LittleEndian<IType, 32> |
ecall
Calls the execution environment, e.g. a Linux system call.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
ecall | emits LittleEndian<IType, 32> |
ebreak
Stops in a debugger.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
ebreak | emits LittleEndian<IType, 32> |
lb
Loads the byte at rs1 + offset into rd, sign-extended.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
lb rd, offset, rs1 | rd: Reg, offset: int, rs1: Reg | emits LittleEndian<IType, 32> |
lh
Loads the 16-bit halfword at rs1 + offset into rd, sign-extended.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
lh rd, offset, rs1 | rd: Reg, offset: int, rs1: Reg | emits LittleEndian<IType, 32> |
lw
Loads the 32-bit word at rs1 + offset into rd.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
lw rd, offset, rs1 | rd: Reg, offset: int, rs1: Reg | emits LittleEndian<IType, 32> |
lbu
Loads the byte at rs1 + offset into rd, zero-extended.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
lbu rd, offset, rs1 | rd: Reg, offset: int, rs1: Reg | emits LittleEndian<IType, 32> |
lhu
Loads the 16-bit halfword at rs1 + offset into rd, zero-extended.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
lhu rd, offset, rs1 | rd: Reg, offset: int, rs1: Reg | emits LittleEndian<IType, 32> |
sb
Stores the low byte of rs2 at rs1 + offset.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
sb rs2, offset, rs1 | rs2: Reg, offset: int, rs1: Reg | emits LittleEndian<SType, 32> |
sh
Stores the low 16 bits of rs2 at rs1 + offset.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
sh rs2, offset, rs1 | rs2: Reg, offset: int, rs1: Reg | emits LittleEndian<SType, 32> |
sw
Stores rs2 at rs1 + offset.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
sw rs2, offset, rs1 | rs2: Reg, offset: int, rs1: Reg | emits LittleEndian<SType, 32> |
beq
Branches to the label target if rs1 == rs2.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
beq rs1, rs2, target | rs1: Reg, rs2: Reg, target: int | emits LittleEndian<BType, 32> |
bne
Branches to the label target if rs1 != rs2.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
bne rs1, rs2, target | rs1: Reg, rs2: Reg, target: int | emits LittleEndian<BType, 32> |
blt
Branches to the label target if rs1 < rs2 as signed numbers.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
blt rs1, rs2, target | rs1: Reg, rs2: Reg, target: int | emits LittleEndian<BType, 32> |
bge
Branches to the label target if rs1 >= rs2 as signed numbers.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
bge rs1, rs2, target | rs1: Reg, rs2: Reg, target: int | emits LittleEndian<BType, 32> |
bltu
Branches to the label target if rs1 < rs2 as unsigned numbers.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
bltu rs1, rs2, target | rs1: Reg, rs2: Reg, target: int | emits LittleEndian<BType, 32> |
bgeu
Branches to the label target if rs1 >= rs2 as unsigned numbers.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
bgeu rs1, rs2, target | rs1: Reg, rs2: Reg, target: int | emits LittleEndian<BType, 32> |
lui
rd = imm << 12: loads a 20-bit upper immediate.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
lui rd, imm | rd: Reg, imm: int | emits LittleEndian<UType, 32> |
auipc
rd = pc + (imm << 12): an address relative to this instruction.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
auipc rd, imm | rd: Reg, imm: int | emits LittleEndian<UType, 32> |
jal
Jumps to the label target and puts the address of the next instruction
in rd. jal ra, f calls f; jal zero, l just jumps.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jal rd, target | rd: Reg, target: int | emits LittleEndian<JType, 32> |
li
Loads the constant imm into rd: addi rd, zero, imm when it fits in
12 signed bits, else lui for the upper 20 bits then, unless they’re
zero, addi for the rest.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
li rd, imm | rd: Reg, imm: int |
la
Loads the address of the label symbol into rd: auipc for the upper
20 bits of the distance, then addi for the rest.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
la rd, symbol | rd: Reg, symbol: int |
ascii
A string literal’s UTF-8 bytes, with no terminator: GNU’s .ascii.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
ascii source | <S>, source: S |
Types
Reg
type Reg = bits<5>
A register number, 0 to 31.
Opcode
type Opcode = bits<7>
The 7-bit major opcode, which picks the format and family.
Funct3
type Funct3 = bits<3>
The 3-bit funct3 field, which picks the instruction within a family.
Funct7
type Funct7 = bits<7>
The 7-bit funct7 field of an R-type instruction.
Bit1
type Bit1 = bits<1>
A one-bit field.
Imm4
type Imm4 = bits<4>
A 4-bit immediate field.
Imm5
type Imm5 = bits<5>
A 5-bit immediate field.
Imm6
type Imm6 = bits<6>
A 6-bit immediate field.
Imm7
type Imm7 = bits<7>
A 7-bit immediate field.
Imm8
type Imm8 = bits<8>
An 8-bit immediate field.
Imm10
type Imm10 = bits<10>
A 10-bit immediate field.
Imm12
type Imm12 = bits<12>
A 12-bit immediate field.
Imm20
type Imm20 = bits<20>
A 20-bit immediate field.
RType
struct RType
The R-type format: register-register operations.
IType
struct IType
The I-type format: register-immediate operations, loads, jalr and
system calls.
SType
struct SType
The S-type format: stores. The 12-bit offset is split around the registers.
BType
struct BType
The B-type format: conditional branches. The offset’s bits are spread
across the word the way the spec lays them out, and resolved by
bitter.
UType
struct UType
The U-type format: a 20-bit upper immediate.
JType
struct JType
The J-type format: jal. Like B-type, its offset’s bits are spread across
the word, and resolved by bitter.
Byte
type Byte = bits<8>
A byte.
Bytes
struct Bytes<const N: int>
N bytes, packed by bitter in order.
Some of its fields are generated by @for or @if.
Constants
| Constant | Type | Value | Description |
|---|---|---|---|
x0 | Reg(0) | A general-purpose register. x0 always reads as zero. | |
x1 | Reg(1) | A general-purpose register. x0 always reads as zero. | |
x2 | Reg(2) | A general-purpose register. x0 always reads as zero. | |
x3 | Reg(3) | A general-purpose register. x0 always reads as zero. | |
x4 | Reg(4) | A general-purpose register. x0 always reads as zero. | |
x5 | Reg(5) | A general-purpose register. x0 always reads as zero. | |
x6 | Reg(6) | A general-purpose register. x0 always reads as zero. | |
x7 | Reg(7) | A general-purpose register. x0 always reads as zero. | |
x8 | Reg(8) | A general-purpose register. x0 always reads as zero. | |
x9 | Reg(9) | A general-purpose register. x0 always reads as zero. | |
x10 | Reg(10) | A general-purpose register. x0 always reads as zero. | |
x11 | Reg(11) | A general-purpose register. x0 always reads as zero. | |
x12 | Reg(12) | A general-purpose register. x0 always reads as zero. | |
x13 | Reg(13) | A general-purpose register. x0 always reads as zero. | |
x14 | Reg(14) | A general-purpose register. x0 always reads as zero. | |
x15 | Reg(15) | A general-purpose register. x0 always reads as zero. | |
x16 | Reg(16) | A general-purpose register. x0 always reads as zero. | |
x17 | Reg(17) | A general-purpose register. x0 always reads as zero. | |
x18 | Reg(18) | A general-purpose register. x0 always reads as zero. | |
x19 | Reg(19) | A general-purpose register. x0 always reads as zero. | |
x20 | Reg(20) | A general-purpose register. x0 always reads as zero. | |
x21 | Reg(21) | A general-purpose register. x0 always reads as zero. | |
x22 | Reg(22) | A general-purpose register. x0 always reads as zero. | |
x23 | Reg(23) | A general-purpose register. x0 always reads as zero. | |
x24 | Reg(24) | A general-purpose register. x0 always reads as zero. | |
x25 | Reg(25) | A general-purpose register. x0 always reads as zero. | |
x26 | Reg(26) | A general-purpose register. x0 always reads as zero. | |
x27 | Reg(27) | A general-purpose register. x0 always reads as zero. | |
x28 | Reg(28) | A general-purpose register. x0 always reads as zero. | |
x29 | Reg(29) | A general-purpose register. x0 always reads as zero. | |
x30 | Reg(30) | A general-purpose register. x0 always reads as zero. | |
x31 | Reg(31) | A general-purpose register. x0 always reads as zero. | |
zero | x0 | x0, which always reads as zero. | |
ra | x1 | x1: the return address. | |
sp | x2 | x2: the stack pointer. | |
gp | x3 | x3: the global pointer. | |
tp | x4 | x4: the thread pointer. | |
t0 | x5 | x5: a temporary. | |
t1 | x6 | x6: a temporary. | |
t2 | x7 | x7: a temporary. | |
s0 | x8 | x8: saved across calls. | |
fp | x8 | x8: the frame pointer, the same register as s0. | |
s1 | x9 | x9: saved across calls. | |
a0 | x10 | x10: a function argument and return value. | |
a1 | x11 | x11: a function argument and return value. | |
a2 | x12 | x12: a function argument. | |
a3 | x13 | x13: a function argument. | |
a4 | x14 | x14: a function argument. | |
a5 | x15 | x15: a function argument. | |
a6 | x16 | x16: a function argument. | |
a7 | x17 | x17: a function argument. | |
s2 | x18 | x18: saved across calls. | |
s3 | x19 | x19: saved across calls. | |
s4 | x20 | x20: saved across calls. | |
s5 | x21 | x21: saved across calls. | |
s6 | x22 | x22: saved across calls. | |
s7 | x23 | x23: saved across calls. | |
s8 | x24 | x24: saved across calls. | |
s9 | x25 | x25: saved across calls. | |
s10 | x26 | x26: saved across calls. | |
s11 | x27 | x27: saved across calls. | |
t3 | x28 | x28: a temporary. | |
t4 | x29 | x29: a temporary. | |
t5 | x30 | x30: a temporary. | |
t6 | x31 | x31: a temporary. |
native
RISC-V RV32I in its standard assembly syntax: add a0, a1, a2,
lw a0, 8(sp), beq a0, zero, done.
from std.riscv.native import *
add a0, a1, a2
lw a0, 8(sp)
33 85 c5 00
03 25 81 00
The instructions themselves, and their registers, come from
std.riscv.impl.
Re-exports std.riscv.impl.
Macros
add
rd = rs1 + rs2.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
add rd, rs1, rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> | std.riscv.impl |
sub
rd = rs1 - rs2.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
sub rd, rs1, rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> | std.riscv.impl |
sll
rd = rs1 << rs2, shifting by the low 5 bits of rs2.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
sll rd, rs1, rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> | std.riscv.impl |
slt
rd = 1 if rs1 < rs2 as signed numbers, else 0.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
slt rd, rs1, rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> | std.riscv.impl |
sltu
rd = 1 if rs1 < rs2 as unsigned numbers, else 0.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
sltu rd, rs1, rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> | std.riscv.impl |
xor
rd = rs1 ^ rs2.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
xor rd, rs1, rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> | std.riscv.impl |
srl
rd = rs1 >> rs2, shifting in zeros, by the low 5 bits of rs2.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
srl rd, rs1, rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> | std.riscv.impl |
sra
rd = rs1 >> rs2, shifting in copies of the sign bit, by the low 5 bits of
rs2.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
sra rd, rs1, rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> | std.riscv.impl |
or
rd = rs1 | rs2.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
or rd, rs1, rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> | std.riscv.impl |
and
rd = rs1 & rs2.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
and rd, rs1, rs2 | rd: Reg, rs1: Reg, rs2: Reg | emits LittleEndian<RType, 32> | std.riscv.impl |
addi
rd = rs1 + imm, with a 12-bit signed imm.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
addi rd, rs1, imm | rd: Reg, rs1: Reg, imm: int | emits LittleEndian<IType, 32> | std.riscv.impl |
slti
rd = 1 if rs1 < imm as signed numbers, else 0.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
slti rd, rs1, imm | rd: Reg, rs1: Reg, imm: int | emits LittleEndian<IType, 32> | std.riscv.impl |
sltiu
rd = 1 if rs1 < imm as unsigned numbers, else 0. imm is
sign-extended first, so sltiu rd, rs1, 1 tests for zero.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
sltiu rd, rs1, imm | rd: Reg, rs1: Reg, imm: int | emits LittleEndian<IType, 32> | std.riscv.impl |
xori
rd = rs1 ^ imm. xori rd, rs1, -1 is bitwise NOT.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
xori rd, rs1, imm | rd: Reg, rs1: Reg, imm: int | emits LittleEndian<IType, 32> | std.riscv.impl |
ori
rd = rs1 | imm.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
ori rd, rs1, imm | rd: Reg, rs1: Reg, imm: int | emits LittleEndian<IType, 32> | std.riscv.impl |
andi
rd = rs1 & imm.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
andi rd, rs1, imm | rd: Reg, rs1: Reg, imm: int | emits LittleEndian<IType, 32> | std.riscv.impl |
slli
rd = rs1 << shamt, for shamt from 0 to 31.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
slli rd, rs1, shamt | rd: Reg, rs1: Reg, shamt: int | emits LittleEndian<IType, 32> | std.riscv.impl |
srli
rd = rs1 >> shamt, shifting in zeros, for shamt from 0 to 31.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
srli rd, rs1, shamt | rd: Reg, rs1: Reg, shamt: int | emits LittleEndian<IType, 32> | std.riscv.impl |
srai
rd = rs1 >> shamt, shifting in copies of the sign bit, for shamt from 0
to 31.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
srai rd, rs1, shamt | rd: Reg, rs1: Reg, shamt: int | emits LittleEndian<IType, 32> | std.riscv.impl |
jalr
Jumps to rs1 + offset (with bit 0 cleared) and puts the address of the
next instruction in rd. jalr zero, ra, 0 returns from a function.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jalr rd, rs1, offset | rd: Reg, rs1: Reg, offset: int | emits LittleEndian<IType, 32> | std.riscv.impl |
ecall
Calls the execution environment, e.g. a Linux system call.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
ecall | emits LittleEndian<IType, 32> | std.riscv.impl |
ebreak
Stops in a debugger.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
ebreak | emits LittleEndian<IType, 32> | std.riscv.impl |
lb
Loads the byte at rs1 + offset into rd, sign-extended.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
lb rd, offset(rs1) | rd: Reg, offset: int, rs1: Reg | emits LittleEndian<IType, 32> | std.riscv.impl |
lh
Loads the 16-bit halfword at rs1 + offset into rd, sign-extended.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
lh rd, offset(rs1) | rd: Reg, offset: int, rs1: Reg | emits LittleEndian<IType, 32> | std.riscv.impl |
lw
Loads the 32-bit word at rs1 + offset into rd.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
lw rd, offset(rs1) | rd: Reg, offset: int, rs1: Reg | emits LittleEndian<IType, 32> | std.riscv.impl |
lbu
Loads the byte at rs1 + offset into rd, zero-extended.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
lbu rd, offset(rs1) | rd: Reg, offset: int, rs1: Reg | emits LittleEndian<IType, 32> | std.riscv.impl |
lhu
Loads the 16-bit halfword at rs1 + offset into rd, zero-extended.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
lhu rd, offset(rs1) | rd: Reg, offset: int, rs1: Reg | emits LittleEndian<IType, 32> | std.riscv.impl |
sb
Stores the low byte of rs2 at rs1 + offset.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
sb rs2, offset(rs1) | rs2: Reg, offset: int, rs1: Reg | emits LittleEndian<SType, 32> | std.riscv.impl |
sh
Stores the low 16 bits of rs2 at rs1 + offset.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
sh rs2, offset(rs1) | rs2: Reg, offset: int, rs1: Reg | emits LittleEndian<SType, 32> | std.riscv.impl |
sw
Stores rs2 at rs1 + offset.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
sw rs2, offset(rs1) | rs2: Reg, offset: int, rs1: Reg | emits LittleEndian<SType, 32> | std.riscv.impl |
beq
Branches to the label target if rs1 == rs2.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
beq rs1, rs2, target | rs1: Reg, rs2: Reg, target: int | emits LittleEndian<BType, 32> | std.riscv.impl |
bne
Branches to the label target if rs1 != rs2.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
bne rs1, rs2, target | rs1: Reg, rs2: Reg, target: int | emits LittleEndian<BType, 32> | std.riscv.impl |
blt
Branches to the label target if rs1 < rs2 as signed numbers.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
blt rs1, rs2, target | rs1: Reg, rs2: Reg, target: int | emits LittleEndian<BType, 32> | std.riscv.impl |
bge
Branches to the label target if rs1 >= rs2 as signed numbers.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
bge rs1, rs2, target | rs1: Reg, rs2: Reg, target: int | emits LittleEndian<BType, 32> | std.riscv.impl |
bltu
Branches to the label target if rs1 < rs2 as unsigned numbers.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
bltu rs1, rs2, target | rs1: Reg, rs2: Reg, target: int | emits LittleEndian<BType, 32> | std.riscv.impl |
bgeu
Branches to the label target if rs1 >= rs2 as unsigned numbers.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
bgeu rs1, rs2, target | rs1: Reg, rs2: Reg, target: int | emits LittleEndian<BType, 32> | std.riscv.impl |
lui
rd = imm << 12: loads a 20-bit upper immediate.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
lui rd, imm | rd: Reg, imm: int | emits LittleEndian<UType, 32> | std.riscv.impl |
auipc
rd = pc + (imm << 12): an address relative to this instruction.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
auipc rd, imm | rd: Reg, imm: int | emits LittleEndian<UType, 32> | std.riscv.impl |
jal
Jumps to the label target and puts the address of the next instruction
in rd. jal ra, f calls f; jal zero, l just jumps.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jal rd, target | rd: Reg, target: int | emits LittleEndian<JType, 32> | std.riscv.impl |
li
Loads the constant imm into rd: addi rd, zero, imm when it fits in
12 signed bits, else lui for the upper 20 bits then, unless they’re
zero, addi for the rest.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
li rd, imm | rd: Reg, imm: int | std.riscv.impl |
la
Loads the address of the label symbol into rd: auipc for the upper
20 bits of the distance, then addi for the rest.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
la rd, symbol | rd: Reg, symbol: int | std.riscv.impl |
ascii
A string literal’s UTF-8 bytes, with no terminator: GNU’s .ascii.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
ascii source | <S>, source: S | std.riscv.impl |
Re-exported
From std.riscv.impl: Reg, x0, x1, x2, x3, x4, x5, x6, x7, x8, x9, x10, x11, x12, x13, x14, x15, x16, x17, x18, x19, x20, x21, x22, x23, x24, x25, x26, x27, x28, x29, x30, x31, zero, ra, sp, gp, tp, t0, t1, t2, s0, fp, s1, a0, a1, a2, a3, a4, a5, a6, a7, s2, s3, s4, s5, s6, s7, s8, s9, s10, s11, t3, t4, t5, t6, Opcode, Funct3, Funct7, Bit1, Imm4, Imm5, Imm6, Imm7, Imm8, Imm10, Imm12, Imm20, RType, IType, SType, BType, UType, JType, Byte, Bytes.
string
std › string
Strings packed into one integer. A string literal like "hi" is a
struct of code points; string_from_struct encodes it as UTF-8 bytes in a
single int, with its length in bytes.
from std.string import *
from std.binary import Endian
macro show(value: int) {
@emit value
}
macro demo() {
const s = string_from_struct("héllo")
show s.len
show utf8_codepoint_count(s as Utf8String<6, Endian.Big>)
show ascii_upper(string_from_struct("hi") as AsciiString<2, Endian.Big>).value
}
demo
6 5 18505
18505 is 0x4849, the bytes of HI. Converting to AsciiString or
Utf8String with as checks the bytes are valid, and picks their
Endian: whether byte 0 is the most (Big) or least (Little)
significant byte of value. The byte order doesn’t change the UTF-8.
Macros
byte_at
Byte index of len bytes packed into value, counting from the
endian end. index must be within the string.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
byte_at(value, len, index, endian) | value: int, len: int, index: int, endian: Endian | returns int |
utf8_struct_byte_len
How many bytes a string literal (or any struct of code points) takes as UTF-8.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
utf8_struct_byte_len(source) | <S>, source: S | returns int |
string_from_struct
A string literal (or any struct of code points) as a String of its
UTF-8 bytes, first character most significant.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
string_from_struct source | <S>, source: S |
validate_ascii
1 if len bytes packed into value are all ASCII. Otherwise it’s a
compile error.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
validate_ascii(value, len, endian) | value: int, len: int, endian: Endian | returns int |
validate_utf8
1 if len bytes packed into value are valid UTF-8: no stray or
missing continuation bytes, overlong forms, surrogates, or code points
past U+10FFFF. Otherwise it’s a compile error.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
validate_utf8(value, len, endian) | value: int, len: int, endian: Endian | returns int |
ascii_byte_at
Byte index of s.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
ascii_byte_at(s, index) | s: AsciiString, index: int | returns int |
utf8_byte_at
Byte index of s. This is a byte, not a character.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
utf8_byte_at(s, index) | s: Utf8String, index: int | returns int |
ascii_upper
s with a-z made uppercase.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
ascii_upper(s) | s: AsciiString | returns AsciiString |
ascii_lower
s with A-Z made lowercase.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
ascii_lower(s) | s: AsciiString | returns AsciiString |
ascii_title
s with its first byte made uppercase.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
ascii_title(s) | s: AsciiString | returns AsciiString |
utf8_codepoint_count
The number of characters (code points) in s.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
utf8_codepoint_count(s) | s: Utf8String | returns int |
utf8_is_ascii
1 if every byte of s is ASCII, else 0.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
utf8_is_ascii(s) | s: Utf8String | returns int |
Types
String
struct String<const len: int>
len bytes packed into value, first byte most significant.
string_from_struct builds one. Convert it with as to AsciiString or
Utf8String to check its bytes and choose their order.
| Field | Type | Description |
|---|---|---|
value | int | The bytes, packed. |
len | int | The number of bytes. |
AsciiString
struct AsciiString<const len: int, const endian: Endian>
len ASCII bytes (each below 0x80) packed into value, with byte 0 at
the endian end.
| Field | Type | Description |
|---|---|---|
value | int | The bytes, packed. |
len | int | The number of bytes. |
endian | Endian | Which end of value holds byte 0. |
Utf8String
struct Utf8String<const len: int, const endian: Endian>
len bytes of valid UTF-8 packed into value, with byte 0 at the
endian end.
| Field | Type | Description |
|---|---|---|
value | int | The bytes, packed. |
len | int | The number of bytes. |
endian | Endian | Which end of value holds byte 0. |
unsigned
std › unsigned
uint, an int that can’t be negative.
Types
uint
type uint = int
An int that’s at least zero, checked wherever a value becomes one.
wasm
std › wasm
| Name | Summary |
|---|---|
impl | WebAssembly instructions: a representative subset of control, variable and i32 instructions, each emitted as its binary encoding. |
leb128 | LEB128, the variable-length integer encoding WebAssembly uses for every index, count and constant: 7 bits per byte, low bits first, with the top bit set on every byte but the last. |
module | The WebAssembly module container: the header, sections, and their length prefixes. A section starts with its id byte and its length in bytes, which deferred_uleb128 fills in once bitter knows it. |
impl
WebAssembly instructions: a representative subset of control, variable
and i32 instructions, each emitted as its binary encoding.
from std.wasm.impl import *
local_get(0)
i32_const(-1)
i32_add
end
20 00 41 7f 6a 0b
WebAssembly’s dotted names aren’t identifiers here, so i32.const is
i32_const and local.get is local_get. Names that are keywords in
other languages end in _: if_, else_, return_. std.wasm.module
builds the module around them.
Macros
unreachable
Traps immediately.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
unreachable | emits Byte |
nop
Does nothing.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
nop | emits Byte |
block
Starts a block. br to it jumps to its end. blocktype is
EMPTY_BLOCKTYPE or the type of the value it produces.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
block blocktype | blocktype: Byte | emits OpWithImm<1> |
loop
Starts a loop. br to it jumps back to its start.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
loop blocktype | blocktype: Byte | emits OpWithImm<1> |
if_
Starts an if: runs what follows if the i32 it pops isn’t zero, and
the else_ part (if any) otherwise.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
if_ blocktype | blocktype: Byte | emits OpWithImm<1> |
else_
Starts the else part of an if_.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
else_ | emits Byte |
end
Ends a block, loop, if_ or function body.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
end | emits Byte |
br
Branches to the enclosing block depth levels out: br(0) is the
innermost.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
br depth | depth: int | emits OpWithImm<...> |
br_if
Pops an i32, and does br(depth) if it isn’t zero.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
br_if depth | depth: int | emits OpWithImm<...> |
return_
Returns from the current function.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
return_ | emits Byte |
call
Calls function number func_index.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
call func_index | func_index: int | emits OpWithImm<...> |
drop
Pops a value and discards it.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
drop | emits Byte |
local_get
Pushes local variable number index.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
local_get index | index: int | emits OpWithImm<...> |
local_set
Pops a value into local variable number index.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
local_set index | index: int | emits OpWithImm<...> |
local_tee
Stores the top of the stack in local variable number index, without
popping it.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
local_tee index | index: int | emits OpWithImm<...> |
i32_load
Pops an address and pushes the i32 at address + offset. align is
the alignment the address is promised to have, as a power of two:
2, a 4-byte boundary, by default.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
i32_load offset, align | offset: int = 0, align: int = 2 | emits MemOp<...> |
i32_store
Pops an i32 value, then an address, and stores the value at
address + offset. align means what it does for i32_load.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
i32_store offset, align | offset: int = 0, align: int = 2 | emits MemOp<...> |
i32_const
Pushes value as an i32.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
i32_const value | value: int | emits OpWithImm<...> |
i32_eqz
Pops an i32 and pushes 1 if it’s zero, else 0.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
i32_eqz | emits Byte |
i32_lt_s
Pops two i32s and pushes 1 if the first is less than the second as
signed numbers, else 0.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
i32_lt_s | emits Byte |
i32_add
Pops two i32s and pushes their sum.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
i32_add | emits Byte |
i32_sub
Pops two i32s and pushes the first minus the second.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
i32_sub | emits Byte |
i32_mul
Pops two i32s and pushes their product.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
i32_mul | emits Byte |
Types
OpWithImm
struct OpWithImm<const N: int>
An opcode followed by N bytes of immediates.
MemOp
struct MemOp<const N: int>
A load or store: its opcode, then its memory operand, the alignment as a power of two and the offset added to the address.
Constants
| Constant | Type | Value | Description |
|---|---|---|---|
I32 | Byte | Byte(0x7F) | The i32 value type. |
I64 | Byte | Byte(0x7E) | The i64 value type. |
F32 | Byte | Byte(0x7D) | The f32 value type. |
F64 | Byte | Byte(0x7C) | The f64 value type. |
EMPTY_BLOCKTYPE | Byte | Byte(0x40) | The block type of a block, loop or if_ that produces no value. A value type such as I32 means it produces one value of that type. |
leb128
LEB128, the variable-length integer encoding WebAssembly uses for every index, count and constant: 7 bits per byte, low bits first, with the top bit set on every byte but the last.
from std.wasm.leb128 import *
macro leb(value: int) {
@emit uleb128(value)
}
macro sleb(value: int) {
@emit sleb128(value)
}
leb 624485
sleb -123456
e5 8e 26
c0 bb 78
Macros
uleb128_length
How many bytes value, which can’t be negative, takes as unsigned
LEB128.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
uleb128_length(value) | value: int | returns int |
uleb128_padded
value as unsigned LEB128 in exactly n bytes. A larger n than
uleb128_length(value) pads with zero groups, which WebAssembly allows,
to give a field a fixed width.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
uleb128_padded(value, n) | value: int, n: int | returns Bytes<...> |
uleb128
value, which can’t be negative, as unsigned LEB128 in as few bytes as
possible.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
uleb128(value) | value: int | returns Bytes<...> |
sleb128_length
How many bytes value takes as signed LEB128.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
sleb128_length(value) | value: int | returns int |
sleb128_padded
value as signed LEB128 in exactly n bytes, sign-extended to fill any
extra ones.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
sleb128_padded(value, n) | value: int, n: int | returns Bytes<...> |
sleb128
value as signed LEB128 in as few bytes as possible.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
sleb128(value) | value: int | returns Bytes<...> |
Types
Byte
type Byte = bits<8>
A byte.
Bytes
struct Bytes<const N: int>
N bytes, packed by bitter in order: a LEB128 encoding, or a whole
instruction.
Some of its fields are generated by @for or @if.
module
The WebAssembly module container: the header, sections, and their
length prefixes. A section starts with its id byte and its length in
bytes, which deferred_uleb128 fills in once bitter knows it.
from std.wasm.module import *
from std.wasm.impl import *
from std.bitter.deferred import *
header
# The type section: one function type, () -> i32.
byte(0x01)
deferred_uleb128 span(types_start, types_end), 1
types_start:
byte(0x01)
func_type_0_to_1(I32)
types_end:
00 61 73 6d 01 00 00 00
01 05 01 60 00 01 7f
section_header writes a section’s id and size, and the helpers below
write the entries of the type, import, function, memory, export, code
and data sections: enough for a WASI program like
examples/wasm/hello.basm. The table, global, start and element sections
have no helpers yet.
Macros
deferred_uleb128
value, usually span(start, end), as unsigned LEB128 in exactly n
bytes, padded if it turns out to need fewer. n has to be enough for
the final value: one byte holds up to 127.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
deferred_uleb128 value, n | value: Deferred, n: int | emits DeferredLeb128<...> |
byte
One byte, such as a section id.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
byte value | value: int | emits Byte |
header
The 8 bytes every module starts with: \0asm, then version 1.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
header | emits Bytes<8> |
func_type_0_to_1
A function type with no parameters and one result of type result,
such as I32.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
func_type_0_to_1 result | result: Byte | emits Bytes<4> |
u32
value, which can’t be negative, as unsigned LEB128: the spec’s u32,
which every count and index is.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
u32 value | value: int | emits Bytes<...> |
size
The number of bytes from label start to label end, as unsigned
LEB128: the size before a section or a function body. It always takes 5
bytes, since it isn’t known until layout, and 5 hold any size.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
size start, end | start: int, end: int | emits DeferredLeb128<5> |
section_header
A section’s id, such as TYPE_SECTION, and its size. Put the label
start right after it and end after the section’s contents.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
section_header id, start, end | id: int, start: int, end: int |
params
| Syntax | Parameters | Result | Description |
|---|---|---|---|
params() | returns Bytes<1> | A function’s parameter types: params(I32, I32). Up to four. | |
params(a) | a: Byte | returns Bytes<2> | |
params(a, b) | a: Byte, b: Byte | returns Bytes<3> | |
params(a, b, c) | a: Byte, b: Byte, c: Byte | returns Bytes<4> | |
params(a, b, c, d) | a: Byte, b: Byte, c: Byte, d: Byte | returns Bytes<5> |
results
| Syntax | Parameters | Result | Description |
|---|---|---|---|
results() | returns Bytes<1> | A function’s result types: results(I32), or results() for none. | |
results(a) | a: Byte | returns Bytes<2> |
func_type
A function type: func_type params(I32), results(I32).
| Syntax | Parameters | Result | Description |
|---|---|---|---|
func_type parameters, returned | parameters: Bytes<...>, returned: Bytes<...> |
name
A name: its length in bytes, then its UTF-8 bytes.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
name source | <S>, source: S |
import_func
Imports the function item_name from the module module_name, with the
type numbered type_index in the type section. Imported functions are
numbered before the module’s own.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
import_func module_name, item_name, type_index | <S, T>, module_name: S, item_name: T, type_index: int |
memory
A memory of pages 64 KiB pages, with no maximum.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
memory pages | pages: int |
export_func
Exports function number index as export_name.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
export_func export_name, index | <S>, export_name: S, index: int |
export_memory
Exports memory number index as export_name.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
export_memory export_name, index | <S>, export_name: S, index: int |
data_segment
A data segment that copies the string source’s UTF-8 bytes into
memory 0 at address offset when the module starts.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
data_segment offset, source | <S>, offset: int, source: S |
Types
Leb128Group
struct Leb128Group
One byte of a LEB128 value bitter works out later.
| Field | Type | Description |
|---|---|---|
continuation | bool | Set on every byte but the last. |
payload | Positioned<7> | Seven bits of the value. |
DeferredLeb128
struct DeferredLeb128<const N: int>
N bytes of a LEB128 value bitter works out later.
Some of its fields are generated by @for or @if.
Constants
| Constant | Type | Value | Description |
|---|---|---|---|
CUSTOM_SECTION | 0 | The id of the custom section. | |
TYPE_SECTION | 1 | The id of the type section: function signatures. | |
IMPORT_SECTION | 2 | The id of the import section. | |
FUNCTION_SECTION | 3 | The id of the function section: each defined function’s type. | |
TABLE_SECTION | 4 | The id of the table section. | |
MEMORY_SECTION | 5 | The id of the memory section. | |
GLOBAL_SECTION | 6 | The id of the global section. | |
EXPORT_SECTION | 7 | The id of the export section. | |
START_SECTION | 8 | The id of the start section. | |
ELEMENT_SECTION | 9 | The id of the element section. | |
CODE_SECTION | 10 | The id of the code section: each defined function’s body. | |
DATA_SECTION | 11 | The id of the data section. |
x86_64
std › x86_64
| Name | Summary |
|---|---|
att | AT&T-syntax x86-64 assembly, as GNU as reads it: source first, destination second, % before registers and $ before immediates. |
impl | The x86-64 base instruction set: general-purpose registers, memory operands, and the core integer instructions, each emitted as its machine code. |
intel | Intel-syntax x86-64 assembly: destination first, registers by name, and memory operands in brackets. |
nasm | NASM-flavored x86-64 assembly: everything std.x86_64.intel provides, plus the NASM spellings it doesn’t have. |
att
AT&T-syntax x86-64 assembly, as GNU as reads it: source first,
destination second, % before registers and $ before immediates.
from std.x86_64.att import *
mov %rbx, %rax
mov 8(%rsp), %rax
add $1, %rax
48 89 d8
48 8b 44 24 08
48 81 c0 01 00 00 00
Memory operands are (%base), disp(%base),
disp(%base,%index,scale) and disp(%rip). Only 64-bit operands have
AT&T spellings, and mnemonics take no size suffix (mov, not movq).
Every instruction of std.x86_64.impl stays available in its explicit
form too: mov rax, rbx, 0 for a 32-bit move.
Re-exports std.x86_64.impl.
Macros
reg_field
The low 3 bits of r’s number: the part a ModRM or SIB field, or an
opcode, holds.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
reg_field(r) | r: Reg | returns int | std.x86_64.impl |
reg_ext
Bit 3 of r’s number, which goes in the REX prefix: 1 for r8 to
r15.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
reg_ext(r) | r: Reg | returns int | std.x86_64.impl |
rex_byte
A REX prefix, 0100WRXB. w selects 64-bit operands; r, x and b
are bit 3 of the ModRM.reg, SIB.index and ModRM.rm (or SIB.base, or
opcode) register numbers.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rex_byte(w, r, x, b) | w: int, r: int, x: int, b: int | returns Byte | std.x86_64.impl |
modrm_byte
A ModRM byte: mod (2 bits), reg (3) and rm (3).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
modrm_byte(mod, reg, rm) | mod: int, reg: int, rm: int | returns Byte | std.x86_64.impl |
sib_byte
A SIB byte: scale (2 bits, log2 of the index’s multiplier), index
(3) and base (3).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
sib_byte(scale, index, base) | scale: int, index: int, base: int | returns Byte | std.x86_64.impl |
byte_of
Byte index of value, counting from the least significant, byte 0.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
byte_of(value, index) | value: int, index: int | returns Byte | std.x86_64.impl |
mov
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
mov rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | rd = rs. | std.x86_64.impl |
mov rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = imm. With w = 1 it takes a full 64-bit imm. | std.x86_64.impl |
mov dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | Stores rs at dst. | std.x86_64.impl |
mov rd, src, w | rd: Reg, src: MemOperand, w: int | emits Bytes<...> | Loads the value at src into rd. | std.x86_64.impl |
mov dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | Stores imm at dst: 32 bits, sign-extended to 64 when w is 1. | std.x86_64.impl |
mov rd, src, w | rd: Reg, src: RipLabel, w: int | Loads the value at the label src into rd. | std.x86_64.impl | |
mov dst, rs, w | dst: RipLabel, rs: Reg, w: int | Stores rs at the label dst. | std.x86_64.impl | |
mov %rs, %rd | rs: Reg, rd: Reg | rd = rs. | ||
mov $imm, %rd | imm: int, rd: Reg | rd = imm. |
add
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
add rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | rd = rd + rs. | std.x86_64.impl |
add rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = rd + imm, with a 32-bit imm sign-extended to 64 bits. | std.x86_64.impl |
add dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | [dst] = [dst] + rs. | std.x86_64.impl |
add dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | [dst] = [dst] + imm, with a 32-bit imm. | std.x86_64.impl |
add dst, rs, w | dst: RipLabel, rs: Reg, w: int | [dst] = [dst] + rs, where dst is a label. | std.x86_64.impl | |
add %rs, %rd | rs: Reg, rd: Reg | rd = rd + rs. | ||
add $imm, %rd | imm: int, rd: Reg | rd = rd + imm. |
or
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
or rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | rd = rd | rs. | std.x86_64.impl |
or rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = rd | imm, with a 32-bit imm sign-extended to 64 bits. | std.x86_64.impl |
or dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | [dst] = [dst] | rs. | std.x86_64.impl |
or dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | [dst] = [dst] | imm, with a 32-bit imm. | std.x86_64.impl |
or dst, rs, w | dst: RipLabel, rs: Reg, w: int | [dst] = [dst] | rs, where dst is a label. | std.x86_64.impl | |
or %rs, %rd | rs: Reg, rd: Reg | rd = rd | rs. | ||
or $imm, %rd | imm: int, rd: Reg | rd = rd | imm. |
and
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
and rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | rd = rd & rs. | std.x86_64.impl |
and rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = rd & imm, with a 32-bit imm sign-extended to 64 bits. | std.x86_64.impl |
and dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | [dst] = [dst] & rs. | std.x86_64.impl |
and dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | [dst] = [dst] & imm, with a 32-bit imm. | std.x86_64.impl |
and dst, rs, w | dst: RipLabel, rs: Reg, w: int | [dst] = [dst] & rs, where dst is a label. | std.x86_64.impl | |
and %rs, %rd | rs: Reg, rd: Reg | rd = rd & rs. | ||
and $imm, %rd | imm: int, rd: Reg | rd = rd & imm. |
sub
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
sub rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | rd = rd - rs. | std.x86_64.impl |
sub rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = rd - imm, with a 32-bit imm sign-extended to 64 bits. | std.x86_64.impl |
sub dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | [dst] = [dst] - rs. | std.x86_64.impl |
sub dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | [dst] = [dst] - imm, with a 32-bit imm. | std.x86_64.impl |
sub dst, rs, w | dst: RipLabel, rs: Reg, w: int | [dst] = [dst] - rs, where dst is a label. | std.x86_64.impl | |
sub %rs, %rd | rs: Reg, rd: Reg | rd = rd - rs. | ||
sub $imm, %rd | imm: int, rd: Reg | rd = rd - imm. |
xor
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
xor rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | rd = rd ^ rs. | std.x86_64.impl |
xor rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = rd ^ imm, with a 32-bit imm sign-extended to 64 bits. | std.x86_64.impl |
xor dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | [dst] = [dst] ^ rs. | std.x86_64.impl |
xor dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | [dst] = [dst] ^ imm, with a 32-bit imm. | std.x86_64.impl |
xor dst, rs, w | dst: RipLabel, rs: Reg, w: int | [dst] = [dst] ^ rs, where dst is a label. | std.x86_64.impl | |
xor %rs, %rd | rs: Reg, rd: Reg | rd = rd ^ rs. | ||
xor $imm, %rd | imm: int, rd: Reg | rd = rd ^ imm. |
cmp
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
cmp rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | Sets the flags from rd - rs, without storing it. | std.x86_64.impl |
cmp rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | Sets the flags from rd - imm, without storing it. | std.x86_64.impl |
cmp dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | Sets the flags from [dst] - rs, without storing it. | std.x86_64.impl |
cmp dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | Sets the flags from [dst] - imm, without storing it. | std.x86_64.impl |
cmp dst, rs, w | dst: RipLabel, rs: Reg, w: int | Sets the flags from [dst] - rs, where dst is a label, without storing it. | std.x86_64.impl | |
cmp %rs, %rd | rs: Reg, rd: Reg | Sets the flags from rd - rs, without storing it. | ||
cmp $imm, %rd | imm: int, rd: Reg | Sets the flags from rd - imm, without storing it. |
test
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
test rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | Sets the flags from rd & rs, without storing it. | std.x86_64.impl |
test rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | Sets the flags from rd & imm, without storing it. | std.x86_64.impl |
test dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | Sets the flags from [dst] & rs, without storing it. | std.x86_64.impl |
test dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | Sets the flags from [dst] & imm, without storing it. | std.x86_64.impl |
test dst, rs, w | dst: RipLabel, rs: Reg, w: int | Sets the flags from [dst] & rs, where dst is a label, without storing it. | std.x86_64.impl | |
test %rs, %rd | rs: Reg, rd: Reg | Sets the flags from rd & rs, without storing it. | ||
test $imm, %rd | imm: int, rd: Reg | Sets the flags from rd & imm, without storing it. |
Mem
The memory at [base + disp]. A displacement from -128 to 127 takes one
byte; any other takes four.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
Mem(base, disp) | base: Reg, disp: int | returns MemOperand | std.x86_64.impl |
MemIndexed
The memory at [base + index * scale + disp]. scale is 1, 2, 4 or 8,
and index can be any register but rsp.
from std.x86_64.impl import *
mov rax, MemIndexed(rbx, r12, 4, 0), 1
4a 8b 04 a3
from std.x86_64.impl import *
mov rax, MemIndexed(rbx, rsp, 4, 0), 1
rsp cannot be a SIB index register
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
MemIndexed(base, index, scale, disp) | base: Reg, index: Reg, scale: int, disp: int | returns MemOperand | std.x86_64.impl |
MemRipRelative
The memory at [rip + disp]: disp bytes past the end of the
instruction. To address a label, use a RipLabel instead.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
MemRipRelative(disp) | disp: int | returns MemOperand | std.x86_64.impl |
jmp
Jumps to the label target.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jmp target | target: int | emits Rel32Instr | std.x86_64.impl |
call
Pushes the address of the next instruction and jumps to the label
target.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
call target | target: int | emits Rel32Instr | std.x86_64.impl |
je
Jumps to the label target if equal (ZF = 1).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
je target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jne
Jumps to the label target if not equal (ZF = 0).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jne target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jb
Jumps to the label target if below, unsigned (CF = 1).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jb target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jae
Jumps to the label target if above or equal, unsigned (CF = 0).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jae target | target: int | emits Rel32Instr2 | std.x86_64.impl |
ja
Jumps to the label target if above, unsigned.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
ja target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jbe
Jumps to the label target if below or equal, unsigned.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jbe target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jl
Jumps to the label target if less, signed.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jl target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jge
Jumps to the label target if greater or equal, signed.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jge target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jle
Jumps to the label target if less or equal, signed.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jle target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jg
Jumps to the label target if greater, signed.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jg target | target: int | emits Rel32Instr2 | std.x86_64.impl |
js
Jumps to the label target if the result was negative (SF = 1).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
js target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jns
Jumps to the label target if the result wasn’t negative (SF = 0).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jns target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jo
Jumps to the label target on signed overflow (OF = 1).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jo target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jno
Jumps to the label target without signed overflow (OF = 0).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jno target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jp
Jumps to the label target if the parity flag is set (PF = 1).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jp target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnp
Jumps to the label target if the parity flag is clear (PF = 0).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnp target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jz
je, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jz target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnz
jne, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnz target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jc
jb, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jc target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnae
jb, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnae target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnc
jae, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnc target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnb
jae, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnb target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnbe
ja, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnbe target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jna
jbe, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jna target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnge
jl, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnge target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnl
jge, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnl target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jng
jle, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jng target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnle
jg, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnle target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jpe
jp, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jpe target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jpo
jnp, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jpo target | target: int | emits Rel32Instr2 | std.x86_64.impl |
ret
Returns: pops an address and jumps to it.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
ret | emits Byte | std.x86_64.impl |
shl
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
shl rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = rd << imm. | std.x86_64.impl |
shl $imm, %rd | imm: int, rd: Reg | rd = rd << imm. |
shl_cl
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
shl_cl(rd, w) | rd: Reg, w: int | emits Bytes<...> | rd = rd << cl. | std.x86_64.impl |
shl %cl, %rd | rd: Reg | rd = rd << cl. |
shr
rd = rd >> imm, shifting in zeros.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
shr rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | std.x86_64.impl |
shr_cl
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
shr_cl(rd, w) | rd: Reg, w: int | emits Bytes<...> | rd = rd >> cl, shifting in zeros. | std.x86_64.impl |
shr %cl, %rd | rd: Reg | rd = rd >> cl, shifting in zeros. |
sar
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
sar rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = rd >> imm, shifting in copies of the sign bit. | std.x86_64.impl |
sar $imm, %rd | imm: int, rd: Reg | rd = rd >> imm, shifting in copies of the sign bit. |
sar_cl
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
sar_cl(rd, w) | rd: Reg, w: int | emits Bytes<...> | rd = rd >> cl, shifting in copies of the sign bit. | std.x86_64.impl |
sar %cl, %rd | rd: Reg | rd = rd >> cl, shifting in copies of the sign bit. |
lea
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
lea rd, src, w | rd: Reg, src: MemOperand, w: int | emits Bytes<...> | Loads the address src stands for into rd, without reading memory. | std.x86_64.impl |
lea rd, src, w | rd: Reg, src: RipLabel, w: int | Loads the address of the label src into rd. | std.x86_64.impl |
rip_label_instr
An instruction with opcode opcode whose memory operand is the label
addr, and whose other operand is r. The dialects’ [rel label]
forms are built on it.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rip_label_instr opcode, addr, r, w | opcode: int, addr: RipLabel, r: Reg, w: int | std.x86_64.impl |
push
Pushes the 64-bit rd onto the stack.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
push %rd | rd: Reg | emits Bytes<...> | std.x86_64.impl |
pop
Pops 64 bits off the stack into rd.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
pop %rd | rd: Reg | emits Bytes<...> | std.x86_64.impl |
syscall
Calls the operating system. On Linux, rax holds the call number and
rdi, rsi, rdx, … its arguments.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
syscall | emits Bytes<2> | std.x86_64.impl |
assert_valid_reg
Fails to compile unless r is a register number, 0 to 15. The
instructions here use it to check their operands.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
assert_valid_reg r | r: Reg |
mov_load_base
Loads the value at [base] into rd.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
mov (%base), %rd | rd: Reg, base: Reg |
mov_load_base_disp
Loads the value at [base + disp] into rd.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
mov disp(%base), %rd | rd: Reg, disp: int, base: Reg |
mov_load_indexed
Loads the value at [base + index * scale + disp] into rd.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
mov disp(%base,%index,scale), %rd | rd: Reg, disp: int, base: Reg, index: Reg, scale: int |
mov_load_rip
Loads the value at [rip + disp], disp bytes past the end of the
instruction, into rd.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
mov disp(%rip), %rd | rd: Reg, disp: int |
mov_store_base
Stores rs at [base].
| Syntax | Parameters | Result | Description |
|---|---|---|---|
mov %rs, (%base) | rs: Reg, base: Reg |
mov_store_base_disp
Stores rs at [base + disp].
| Syntax | Parameters | Result | Description |
|---|---|---|---|
mov %rs, disp(%base) | rs: Reg, disp: int, base: Reg |
mov_store_indexed
Stores rs at [base + index * scale + disp].
| Syntax | Parameters | Result | Description |
|---|---|---|---|
mov %rs, disp(%base,%index,scale) | rs: Reg, disp: int, base: Reg, index: Reg, scale: int |
mov_store_rip
Stores rs at [rip + disp], disp bytes past the end of the
instruction.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
mov %rs, disp(%rip) | rs: Reg, disp: int |
lea_base
Loads the address [base] into rd, without reading memory.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
lea (%base), %rd | rd: Reg, base: Reg |
lea_base_disp
Loads the address [base + disp] into rd, without reading memory.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
lea disp(%base), %rd | rd: Reg, disp: int, base: Reg |
lea_indexed
Loads the address [base + index * scale + disp] into rd, without reading
memory.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
lea disp(%base,%index,scale), %rd | rd: Reg, disp: int, base: Reg, index: Reg, scale: int |
lea_rip
Loads the address [rip + disp], disp bytes past the end of the
instruction, into rd.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
lea disp(%rip), %rd | rd: Reg, disp: int |
shr_imm
rd = rd >> imm, shifting in zeros.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
shr $imm, %rd | imm: int, rd: Reg |
Re-exported
From std.x86_64.impl: Reg, r0, r1, r2, r3, r4, r5, r6, r7, r8, r9, r10, r11, r12, r13, r14, r15, rax, rcx, rdx, rbx, rsp, rbp, rsi, rdi, Reg32, eax, ecx, edx, ebx, esp, ebp, esi, edi, r8d, r9d, r10d, r11d, r12d, r13d, r14d, r15d, Byte, Bytes, MemBase, MemSib, MemRip, MemOperand, Rel32Instr, Rel32Instr2, RipLabel, RipRelInstr, RipRelInstrRex.
impl
The x86-64 base instruction set: general-purpose registers, memory operands, and the core integer instructions, each emitted as its machine code.
Import a dialect rather than this module: std.x86_64.intel,
std.x86_64.nasm or std.x86_64.att. Here every instruction takes its
operands in order, destination first, plus a final w: 1 for 64-bit
operands and 0 for 32-bit. The dialects pick w from the register’s
name instead (rax or eax).
from std.x86_64.impl import *
mov rax, rbx, 1
mov rax, Mem(rsp, 8), 1
loop:
sub rcx, 1, 1
jne loop
48 89 d8
48 8b 44 24 08
48 81 e9 01 00 00 00
0f 85 f3 ff ff ff
Jumps and calls always use a 32-bit offset, worked out by bitter once
the program is laid out; there’s no automatic choice of the short form.
Macros
reg_field
The low 3 bits of r’s number: the part a ModRM or SIB field, or an
opcode, holds.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
reg_field(r) | r: Reg | returns int |
reg_ext
Bit 3 of r’s number, which goes in the REX prefix: 1 for r8 to
r15.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
reg_ext(r) | r: Reg | returns int |
rex_byte
A REX prefix, 0100WRXB. w selects 64-bit operands; r, x and b
are bit 3 of the ModRM.reg, SIB.index and ModRM.rm (or SIB.base, or
opcode) register numbers.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
rex_byte(w, r, x, b) | w: int, r: int, x: int, b: int | returns Byte |
modrm_byte
A ModRM byte: mod (2 bits), reg (3) and rm (3).
| Syntax | Parameters | Result | Description |
|---|---|---|---|
modrm_byte(mod, reg, rm) | mod: int, reg: int, rm: int | returns Byte |
sib_byte
A SIB byte: scale (2 bits, log2 of the index’s multiplier), index
(3) and base (3).
| Syntax | Parameters | Result | Description |
|---|---|---|---|
sib_byte(scale, index, base) | scale: int, index: int, base: int | returns Byte |
byte_of
Byte index of value, counting from the least significant, byte 0.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
byte_of(value, index) | value: int, index: int | returns Byte |
mov
| Syntax | Parameters | Result | Description |
|---|---|---|---|
mov rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | rd = rs. |
mov rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = imm. With w = 1 it takes a full 64-bit imm. |
mov dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | Stores rs at dst. |
mov rd, src, w | rd: Reg, src: MemOperand, w: int | emits Bytes<...> | Loads the value at src into rd. |
mov dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | Stores imm at dst: 32 bits, sign-extended to 64 when w is 1. |
mov rd, src, w | rd: Reg, src: RipLabel, w: int | Loads the value at the label src into rd. | |
mov dst, rs, w | dst: RipLabel, rs: Reg, w: int | Stores rs at the label dst. |
add
| Syntax | Parameters | Result | Description |
|---|---|---|---|
add rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | rd = rd + rs. |
add rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = rd + imm, with a 32-bit imm sign-extended to 64 bits. |
add dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | [dst] = [dst] + rs. |
add dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | [dst] = [dst] + imm, with a 32-bit imm. |
add dst, rs, w | dst: RipLabel, rs: Reg, w: int | [dst] = [dst] + rs, where dst is a label. |
or
| Syntax | Parameters | Result | Description |
|---|---|---|---|
or rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | rd = rd | rs. |
or rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = rd | imm, with a 32-bit imm sign-extended to 64 bits. |
or dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | [dst] = [dst] | rs. |
or dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | [dst] = [dst] | imm, with a 32-bit imm. |
or dst, rs, w | dst: RipLabel, rs: Reg, w: int | [dst] = [dst] | rs, where dst is a label. |
and
| Syntax | Parameters | Result | Description |
|---|---|---|---|
and rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | rd = rd & rs. |
and rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = rd & imm, with a 32-bit imm sign-extended to 64 bits. |
and dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | [dst] = [dst] & rs. |
and dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | [dst] = [dst] & imm, with a 32-bit imm. |
and dst, rs, w | dst: RipLabel, rs: Reg, w: int | [dst] = [dst] & rs, where dst is a label. |
sub
| Syntax | Parameters | Result | Description |
|---|---|---|---|
sub rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | rd = rd - rs. |
sub rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = rd - imm, with a 32-bit imm sign-extended to 64 bits. |
sub dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | [dst] = [dst] - rs. |
sub dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | [dst] = [dst] - imm, with a 32-bit imm. |
sub dst, rs, w | dst: RipLabel, rs: Reg, w: int | [dst] = [dst] - rs, where dst is a label. |
xor
| Syntax | Parameters | Result | Description |
|---|---|---|---|
xor rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | rd = rd ^ rs. |
xor rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = rd ^ imm, with a 32-bit imm sign-extended to 64 bits. |
xor dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | [dst] = [dst] ^ rs. |
xor dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | [dst] = [dst] ^ imm, with a 32-bit imm. |
xor dst, rs, w | dst: RipLabel, rs: Reg, w: int | [dst] = [dst] ^ rs, where dst is a label. |
cmp
| Syntax | Parameters | Result | Description |
|---|---|---|---|
cmp rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | Sets the flags from rd - rs, without storing it. |
cmp rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | Sets the flags from rd - imm, without storing it. |
cmp dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | Sets the flags from [dst] - rs, without storing it. |
cmp dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | Sets the flags from [dst] - imm, without storing it. |
cmp dst, rs, w | dst: RipLabel, rs: Reg, w: int | Sets the flags from [dst] - rs, where dst is a label, without storing it. |
test
| Syntax | Parameters | Result | Description |
|---|---|---|---|
test rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | Sets the flags from rd & rs, without storing it. |
test rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | Sets the flags from rd & imm, without storing it. |
test dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | Sets the flags from [dst] & rs, without storing it. |
test dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | Sets the flags from [dst] & imm, without storing it. |
test dst, rs, w | dst: RipLabel, rs: Reg, w: int | Sets the flags from [dst] & rs, where dst is a label, without storing it. |
Mem
The memory at [base + disp]. A displacement from -128 to 127 takes one
byte; any other takes four.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
Mem(base, disp) | base: Reg, disp: int | returns MemOperand |
MemIndexed
The memory at [base + index * scale + disp]. scale is 1, 2, 4 or 8,
and index can be any register but rsp.
from std.x86_64.impl import *
mov rax, MemIndexed(rbx, r12, 4, 0), 1
4a 8b 04 a3
from std.x86_64.impl import *
mov rax, MemIndexed(rbx, rsp, 4, 0), 1
rsp cannot be a SIB index register
| Syntax | Parameters | Result | Description |
|---|---|---|---|
MemIndexed(base, index, scale, disp) | base: Reg, index: Reg, scale: int, disp: int | returns MemOperand |
MemRipRelative
The memory at [rip + disp]: disp bytes past the end of the
instruction. To address a label, use a RipLabel instead.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
MemRipRelative(disp) | disp: int | returns MemOperand |
jmp
Jumps to the label target.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jmp target | target: int | emits Rel32Instr |
call
Pushes the address of the next instruction and jumps to the label
target.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
call target | target: int | emits Rel32Instr |
je
Jumps to the label target if equal (ZF = 1).
| Syntax | Parameters | Result | Description |
|---|---|---|---|
je target | target: int | emits Rel32Instr2 |
jne
Jumps to the label target if not equal (ZF = 0).
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jne target | target: int | emits Rel32Instr2 |
jb
Jumps to the label target if below, unsigned (CF = 1).
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jb target | target: int | emits Rel32Instr2 |
jae
Jumps to the label target if above or equal, unsigned (CF = 0).
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jae target | target: int | emits Rel32Instr2 |
ja
Jumps to the label target if above, unsigned.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
ja target | target: int | emits Rel32Instr2 |
jbe
Jumps to the label target if below or equal, unsigned.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jbe target | target: int | emits Rel32Instr2 |
jl
Jumps to the label target if less, signed.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jl target | target: int | emits Rel32Instr2 |
jge
Jumps to the label target if greater or equal, signed.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jge target | target: int | emits Rel32Instr2 |
jle
Jumps to the label target if less or equal, signed.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jle target | target: int | emits Rel32Instr2 |
jg
Jumps to the label target if greater, signed.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jg target | target: int | emits Rel32Instr2 |
js
Jumps to the label target if the result was negative (SF = 1).
| Syntax | Parameters | Result | Description |
|---|---|---|---|
js target | target: int | emits Rel32Instr2 |
jns
Jumps to the label target if the result wasn’t negative (SF = 0).
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jns target | target: int | emits Rel32Instr2 |
jo
Jumps to the label target on signed overflow (OF = 1).
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jo target | target: int | emits Rel32Instr2 |
jno
Jumps to the label target without signed overflow (OF = 0).
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jno target | target: int | emits Rel32Instr2 |
jp
Jumps to the label target if the parity flag is set (PF = 1).
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jp target | target: int | emits Rel32Instr2 |
jnp
Jumps to the label target if the parity flag is clear (PF = 0).
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jnp target | target: int | emits Rel32Instr2 |
jz
je, under another name.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jz target | target: int | emits Rel32Instr2 |
jnz
jne, under another name.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jnz target | target: int | emits Rel32Instr2 |
jc
jb, under another name.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jc target | target: int | emits Rel32Instr2 |
jnae
jb, under another name.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jnae target | target: int | emits Rel32Instr2 |
jnc
jae, under another name.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jnc target | target: int | emits Rel32Instr2 |
jnb
jae, under another name.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jnb target | target: int | emits Rel32Instr2 |
jnbe
ja, under another name.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jnbe target | target: int | emits Rel32Instr2 |
jna
jbe, under another name.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jna target | target: int | emits Rel32Instr2 |
jnge
jl, under another name.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jnge target | target: int | emits Rel32Instr2 |
jnl
jge, under another name.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jnl target | target: int | emits Rel32Instr2 |
jng
jle, under another name.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jng target | target: int | emits Rel32Instr2 |
jnle
jg, under another name.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jnle target | target: int | emits Rel32Instr2 |
jpe
jp, under another name.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jpe target | target: int | emits Rel32Instr2 |
jpo
jnp, under another name.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
jpo target | target: int | emits Rel32Instr2 |
ret
Returns: pops an address and jumps to it.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
ret | emits Byte |
shl
rd = rd << imm.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
shl rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> |
shl_cl
rd = rd << cl.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
shl_cl rd, w | rd: Reg, w: int | emits Bytes<...> |
shr
rd = rd >> imm, shifting in zeros.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
shr rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> |
shr_cl
rd = rd >> cl, shifting in zeros.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
shr_cl rd, w | rd: Reg, w: int | emits Bytes<...> |
sar
rd = rd >> imm, shifting in copies of the sign bit.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
sar rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> |
sar_cl
rd = rd >> cl, shifting in copies of the sign bit.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
sar_cl rd, w | rd: Reg, w: int | emits Bytes<...> |
lea
| Syntax | Parameters | Result | Description |
|---|---|---|---|
lea rd, src, w | rd: Reg, src: MemOperand, w: int | emits Bytes<...> | Loads the address src stands for into rd, without reading memory. |
lea rd, src, w | rd: Reg, src: RipLabel, w: int | Loads the address of the label src into rd. |
rip_label_instr
An instruction with opcode opcode whose memory operand is the label
addr, and whose other operand is r. The dialects’ [rel label]
forms are built on it.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
rip_label_instr opcode, addr, r, w | opcode: int, addr: RipLabel, r: Reg, w: int |
push
Pushes the 64-bit rd onto the stack.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
push rd | rd: Reg | emits Bytes<...> |
pop
Pops 64 bits off the stack into rd.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
pop rd | rd: Reg | emits Bytes<...> |
syscall
Calls the operating system. On Linux, rax holds the call number and
rdi, rsi, rdx, … its arguments.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
syscall | emits Bytes<2> |
Types
Reg
type Reg = bits<4>
A register number, 0 to 15. As an operand it means the whole 64-bit register.
Reg32
struct Reg32
The low 32 bits of a register. An instruction given one works on 32 bits instead of 64, and writing one zeroes the register’s upper half.
| Field | Type | Description |
|---|---|---|
reg | Reg | The register it’s the low half of. |
Byte
type Byte = bits<8>
A byte.
Bytes
struct Bytes<const N: int>
N bytes, packed by bitter in order: an instruction’s encoding.
Some of its fields are generated by @for or @if.
MemBase
struct MemBase
[base + disp]. Mem builds one.
| Field | Type | Description |
|---|---|---|
base | Reg | The base register. |
disp | int | The displacement added to it. |
MemSib
struct MemSib
[base + index * scale + disp]. MemIndexed builds one.
| Field | Type | Description |
|---|---|---|
base | Reg | The base register. |
index | Reg | The index register. |
scale | int | The index’s multiplier: 1, 2, 4 or 8. |
disp | int | The displacement. |
MemRip
struct MemRip
[rip + disp]. MemRipRelative builds one.
| Field | Type | Description |
|---|---|---|
disp | int | The displacement from the end of the instruction. |
MemOperand
enum MemOperand
A memory operand, built by Mem, MemIndexed or MemRipRelative.
| Variant | Payload | Description |
|---|---|---|
Base | MemBase | [base + disp]. |
Indexed | MemSib | [base + index * scale + disp]. |
RipRelative | MemRip | [rip + disp]. |
Rel32Instr
struct Rel32Instr
A jmp or call: an opcode byte and a 32-bit offset bitter works out.
| Field | Type | Description |
|---|---|---|
opcode | Byte | The opcode. |
rel32 | LittleEndian<Positioned<32>, 32> | The offset from the next instruction to the target. |
Rel32Instr2
struct Rel32Instr2
A conditional jump: two opcode bytes and a 32-bit offset bitter works
out.
| Field | Type | Description |
|---|---|---|
opcode1 | Byte | The first opcode byte, 0x0F. |
opcode2 | Byte | The second opcode byte, which holds the condition. |
rel32 | LittleEndian<Positioned<32>, 32> | The offset from the next instruction to the target. |
RipLabel
struct RipLabel
A label used as a memory operand, addressed relative to the next
instruction: NASM’s [rel label].
| Field | Type | Description |
|---|---|---|
target | int | The label. |
RipRelInstr
struct RipRelInstr
An instruction with a RipLabel operand and no REX prefix.
| Field | Type | Description |
|---|---|---|
opcode | Byte | The opcode. |
modrm | Byte | The ModRM byte, with rm = RIP-relative. |
disp32 | LittleEndian<Positioned<32>, 32> | The displacement from the next instruction to the label. |
RipRelInstrRex
struct RipRelInstrRex
An instruction with a RipLabel operand and a REX prefix.
| Field | Type | Description |
|---|---|---|
rex | Byte | The REX prefix. |
opcode | Byte | The opcode. |
modrm | Byte | The ModRM byte, with rm = RIP-relative. |
disp32 | LittleEndian<Positioned<32>, 32> | The displacement from the next instruction to the label. |
Constants
| Constant | Type | Value | Description |
|---|---|---|---|
r0 | Reg | 0 | A 64-bit general-purpose register. r0 to r7 are better known as rax to rdi. |
r1 | Reg | 1 | A 64-bit general-purpose register. r0 to r7 are better known as rax to rdi. |
r2 | Reg | 2 | A 64-bit general-purpose register. r0 to r7 are better known as rax to rdi. |
r3 | Reg | 3 | A 64-bit general-purpose register. r0 to r7 are better known as rax to rdi. |
r4 | Reg | 4 | A 64-bit general-purpose register. r0 to r7 are better known as rax to rdi. |
r5 | Reg | 5 | A 64-bit general-purpose register. r0 to r7 are better known as rax to rdi. |
r6 | Reg | 6 | A 64-bit general-purpose register. r0 to r7 are better known as rax to rdi. |
r7 | Reg | 7 | A 64-bit general-purpose register. r0 to r7 are better known as rax to rdi. |
r8 | Reg | 8 | A 64-bit general-purpose register. r0 to r7 are better known as rax to rdi. |
r9 | Reg | 9 | A 64-bit general-purpose register. r0 to r7 are better known as rax to rdi. |
r10 | Reg | 10 | A 64-bit general-purpose register. r0 to r7 are better known as rax to rdi. |
r11 | Reg | 11 | A 64-bit general-purpose register. r0 to r7 are better known as rax to rdi. |
r12 | Reg | 12 | A 64-bit general-purpose register. r0 to r7 are better known as rax to rdi. |
r13 | Reg | 13 | A 64-bit general-purpose register. r0 to r7 are better known as rax to rdi. |
r14 | Reg | 14 | A 64-bit general-purpose register. r0 to r7 are better known as rax to rdi. |
r15 | Reg | 15 | A 64-bit general-purpose register. r0 to r7 are better known as rax to rdi. |
rax | r0 | r0, the accumulator: where a function returns its result. | |
rcx | r1 | r1, the counter: a function’s fourth argument. | |
rdx | r2 | r2: a function’s third argument. | |
rbx | r3 | r3, saved across calls. | |
rsp | r4 | r4, the stack pointer. | |
rbp | r5 | r5, the frame pointer, saved across calls. | |
rsi | r6 | r6: a function’s second argument. | |
rdi | r7 | r7: a function’s first argument. | |
eax | Reg32(reg = rax) | The low 32 bits of rax. | |
ecx | Reg32(reg = rcx) | The low 32 bits of rcx. | |
edx | Reg32(reg = rdx) | The low 32 bits of rdx. | |
ebx | Reg32(reg = rbx) | The low 32 bits of rbx. | |
esp | Reg32(reg = rsp) | The low 32 bits of rsp. | |
ebp | Reg32(reg = rbp) | The low 32 bits of rbp. | |
esi | Reg32(reg = rsi) | The low 32 bits of rsi. | |
edi | Reg32(reg = rdi) | The low 32 bits of rdi. | |
r8d | Reg32(reg = r8) | The low 32 bits of r8. | |
r9d | Reg32(reg = r9) | The low 32 bits of r9. | |
r10d | Reg32(reg = r10) | The low 32 bits of r10. | |
r11d | Reg32(reg = r11) | The low 32 bits of r11. | |
r12d | Reg32(reg = r12) | The low 32 bits of r12. | |
r13d | Reg32(reg = r13) | The low 32 bits of r13. | |
r14d | Reg32(reg = r14) | The low 32 bits of r14. | |
r15d | Reg32(reg = r15) | The low 32 bits of r15. |
intel
Intel-syntax x86-64 assembly: destination first, registers by name, and memory operands in brackets.
from std.x86_64.intel import *
mov rax, [rsp+8]
add eax, 1
mov r8, [rbx+rcx*4+0x10]
48 8b 44 24 08
81 c0 01 00 00 00
4c 8b 44 8b 10
The operand size comes from the register’s name: rax, r8 and the
other Regs are 64-bit, and eax, r8d and the other Reg32s are
32-bit. Memory operands are [base], [base+disp],
[base+index*scale+disp] and [rip+disp]. Every instruction of
std.x86_64.impl stays available in its explicit form too, with the
operand size as a final argument: mov rax, rbx, 0.
For NASM’s [rel label] and data directives, use std.x86_64.nasm.
Re-exports std.x86_64.impl.
Macros
reg_field
The low 3 bits of r’s number: the part a ModRM or SIB field, or an
opcode, holds.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
reg_field(r) | r: Reg | returns int | std.x86_64.impl |
reg_ext
Bit 3 of r’s number, which goes in the REX prefix: 1 for r8 to
r15.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
reg_ext(r) | r: Reg | returns int | std.x86_64.impl |
rex_byte
A REX prefix, 0100WRXB. w selects 64-bit operands; r, x and b
are bit 3 of the ModRM.reg, SIB.index and ModRM.rm (or SIB.base, or
opcode) register numbers.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rex_byte(w, r, x, b) | w: int, r: int, x: int, b: int | returns Byte | std.x86_64.impl |
modrm_byte
A ModRM byte: mod (2 bits), reg (3) and rm (3).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
modrm_byte(mod, reg, rm) | mod: int, reg: int, rm: int | returns Byte | std.x86_64.impl |
sib_byte
A SIB byte: scale (2 bits, log2 of the index’s multiplier), index
(3) and base (3).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
sib_byte(scale, index, base) | scale: int, index: int, base: int | returns Byte | std.x86_64.impl |
byte_of
Byte index of value, counting from the least significant, byte 0.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
byte_of(value, index) | value: int, index: int | returns Byte | std.x86_64.impl |
mov
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
mov rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | rd = rs. | std.x86_64.impl |
mov rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = imm. With w = 1 it takes a full 64-bit imm. | std.x86_64.impl |
mov dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | Stores rs at dst. | std.x86_64.impl |
mov rd, src, w | rd: Reg, src: MemOperand, w: int | emits Bytes<...> | Loads the value at src into rd. | std.x86_64.impl |
mov dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | Stores imm at dst: 32 bits, sign-extended to 64 when w is 1. | std.x86_64.impl |
mov rd, src, w | rd: Reg, src: RipLabel, w: int | Loads the value at the label src into rd. | std.x86_64.impl | |
mov dst, rs, w | dst: RipLabel, rs: Reg, w: int | Stores rs at the label dst. | std.x86_64.impl | |
mov rd, rs | rd: Reg, rs: Reg | rd = rs. | ||
mov rd, imm | rd: Reg, imm: int | rd = imm. | ||
mov rd, rs | rd: Reg32, rs: Reg32 | rd = rs, in 32 bits. | ||
mov rd, imm | rd: Reg32, imm: int | rd = imm, in 32 bits. |
add
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
add rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | rd = rd + rs. | std.x86_64.impl |
add rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = rd + imm, with a 32-bit imm sign-extended to 64 bits. | std.x86_64.impl |
add dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | [dst] = [dst] + rs. | std.x86_64.impl |
add dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | [dst] = [dst] + imm, with a 32-bit imm. | std.x86_64.impl |
add dst, rs, w | dst: RipLabel, rs: Reg, w: int | [dst] = [dst] + rs, where dst is a label. | std.x86_64.impl | |
add rd, rs | rd: Reg, rs: Reg | rd = rd + rs. | ||
add rd, imm | rd: Reg, imm: int | rd = rd + imm. | ||
add rd, rs | rd: Reg32, rs: Reg32 | rd = rd + rs, in 32 bits. | ||
add rd, imm | rd: Reg32, imm: int | rd = rd + imm, in 32 bits. |
or
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
or rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | rd = rd | rs. | std.x86_64.impl |
or rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = rd | imm, with a 32-bit imm sign-extended to 64 bits. | std.x86_64.impl |
or dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | [dst] = [dst] | rs. | std.x86_64.impl |
or dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | [dst] = [dst] | imm, with a 32-bit imm. | std.x86_64.impl |
or dst, rs, w | dst: RipLabel, rs: Reg, w: int | [dst] = [dst] | rs, where dst is a label. | std.x86_64.impl | |
or rd, rs | rd: Reg, rs: Reg | rd = rd | rs. | ||
or rd, imm | rd: Reg, imm: int | rd = rd | imm. | ||
or rd, rs | rd: Reg32, rs: Reg32 | rd = rd | rs, in 32 bits. | ||
or rd, imm | rd: Reg32, imm: int | rd = rd | imm, in 32 bits. |
and
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
and rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | rd = rd & rs. | std.x86_64.impl |
and rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = rd & imm, with a 32-bit imm sign-extended to 64 bits. | std.x86_64.impl |
and dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | [dst] = [dst] & rs. | std.x86_64.impl |
and dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | [dst] = [dst] & imm, with a 32-bit imm. | std.x86_64.impl |
and dst, rs, w | dst: RipLabel, rs: Reg, w: int | [dst] = [dst] & rs, where dst is a label. | std.x86_64.impl | |
and rd, rs | rd: Reg, rs: Reg | rd = rd & rs. | ||
and rd, imm | rd: Reg, imm: int | rd = rd & imm. | ||
and rd, rs | rd: Reg32, rs: Reg32 | rd = rd & rs, in 32 bits. | ||
and rd, imm | rd: Reg32, imm: int | rd = rd & imm, in 32 bits. |
sub
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
sub rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | rd = rd - rs. | std.x86_64.impl |
sub rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = rd - imm, with a 32-bit imm sign-extended to 64 bits. | std.x86_64.impl |
sub dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | [dst] = [dst] - rs. | std.x86_64.impl |
sub dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | [dst] = [dst] - imm, with a 32-bit imm. | std.x86_64.impl |
sub dst, rs, w | dst: RipLabel, rs: Reg, w: int | [dst] = [dst] - rs, where dst is a label. | std.x86_64.impl | |
sub rd, rs | rd: Reg, rs: Reg | rd = rd - rs. | ||
sub rd, imm | rd: Reg, imm: int | rd = rd - imm. | ||
sub rd, rs | rd: Reg32, rs: Reg32 | rd = rd - rs, in 32 bits. | ||
sub rd, imm | rd: Reg32, imm: int | rd = rd - imm, in 32 bits. |
xor
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
xor rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | rd = rd ^ rs. | std.x86_64.impl |
xor rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = rd ^ imm, with a 32-bit imm sign-extended to 64 bits. | std.x86_64.impl |
xor dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | [dst] = [dst] ^ rs. | std.x86_64.impl |
xor dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | [dst] = [dst] ^ imm, with a 32-bit imm. | std.x86_64.impl |
xor dst, rs, w | dst: RipLabel, rs: Reg, w: int | [dst] = [dst] ^ rs, where dst is a label. | std.x86_64.impl | |
xor rd, rs | rd: Reg, rs: Reg | rd = rd ^ rs. | ||
xor rd, imm | rd: Reg, imm: int | rd = rd ^ imm. | ||
xor rd, rs | rd: Reg32, rs: Reg32 | rd = rd ^ rs, in 32 bits. | ||
xor rd, imm | rd: Reg32, imm: int | rd = rd ^ imm, in 32 bits. |
cmp
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
cmp rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | Sets the flags from rd - rs, without storing it. | std.x86_64.impl |
cmp rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | Sets the flags from rd - imm, without storing it. | std.x86_64.impl |
cmp dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | Sets the flags from [dst] - rs, without storing it. | std.x86_64.impl |
cmp dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | Sets the flags from [dst] - imm, without storing it. | std.x86_64.impl |
cmp dst, rs, w | dst: RipLabel, rs: Reg, w: int | Sets the flags from [dst] - rs, where dst is a label, without storing it. | std.x86_64.impl | |
cmp rd, rs | rd: Reg, rs: Reg | Sets the flags from rd - rs, without storing it. | ||
cmp rd, imm | rd: Reg, imm: int | Sets the flags from rd - imm, without storing it. | ||
cmp rd, rs | rd: Reg32, rs: Reg32 | Sets the flags from rd - rs, in 32 bits, without storing it. | ||
cmp rd, imm | rd: Reg32, imm: int | Sets the flags from rd - imm, in 32 bits, without storing it. |
test
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
test rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | Sets the flags from rd & rs, without storing it. | std.x86_64.impl |
test rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | Sets the flags from rd & imm, without storing it. | std.x86_64.impl |
test dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | Sets the flags from [dst] & rs, without storing it. | std.x86_64.impl |
test dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | Sets the flags from [dst] & imm, without storing it. | std.x86_64.impl |
test dst, rs, w | dst: RipLabel, rs: Reg, w: int | Sets the flags from [dst] & rs, where dst is a label, without storing it. | std.x86_64.impl | |
test rd, rs | rd: Reg, rs: Reg | Sets the flags from rd & rs, without storing it. | ||
test rd, imm | rd: Reg, imm: int | Sets the flags from rd & imm, without storing it. | ||
test rd, rs | rd: Reg32, rs: Reg32 | Sets the flags from rd & rs, in 32 bits, without storing it. | ||
test rd, imm | rd: Reg32, imm: int | Sets the flags from rd & imm, in 32 bits, without storing it. |
Mem
The memory at [base + disp]. A displacement from -128 to 127 takes one
byte; any other takes four.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
Mem(base, disp) | base: Reg, disp: int | returns MemOperand | std.x86_64.impl |
MemIndexed
The memory at [base + index * scale + disp]. scale is 1, 2, 4 or 8,
and index can be any register but rsp.
from std.x86_64.impl import *
mov rax, MemIndexed(rbx, r12, 4, 0), 1
4a 8b 04 a3
from std.x86_64.impl import *
mov rax, MemIndexed(rbx, rsp, 4, 0), 1
rsp cannot be a SIB index register
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
MemIndexed(base, index, scale, disp) | base: Reg, index: Reg, scale: int, disp: int | returns MemOperand | std.x86_64.impl |
MemRipRelative
The memory at [rip + disp]: disp bytes past the end of the
instruction. To address a label, use a RipLabel instead.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
MemRipRelative(disp) | disp: int | returns MemOperand | std.x86_64.impl |
jmp
Jumps to the label target.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jmp target | target: int | emits Rel32Instr | std.x86_64.impl |
call
Pushes the address of the next instruction and jumps to the label
target.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
call target | target: int | emits Rel32Instr | std.x86_64.impl |
je
Jumps to the label target if equal (ZF = 1).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
je target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jne
Jumps to the label target if not equal (ZF = 0).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jne target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jb
Jumps to the label target if below, unsigned (CF = 1).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jb target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jae
Jumps to the label target if above or equal, unsigned (CF = 0).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jae target | target: int | emits Rel32Instr2 | std.x86_64.impl |
ja
Jumps to the label target if above, unsigned.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
ja target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jbe
Jumps to the label target if below or equal, unsigned.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jbe target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jl
Jumps to the label target if less, signed.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jl target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jge
Jumps to the label target if greater or equal, signed.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jge target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jle
Jumps to the label target if less or equal, signed.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jle target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jg
Jumps to the label target if greater, signed.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jg target | target: int | emits Rel32Instr2 | std.x86_64.impl |
js
Jumps to the label target if the result was negative (SF = 1).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
js target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jns
Jumps to the label target if the result wasn’t negative (SF = 0).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jns target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jo
Jumps to the label target on signed overflow (OF = 1).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jo target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jno
Jumps to the label target without signed overflow (OF = 0).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jno target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jp
Jumps to the label target if the parity flag is set (PF = 1).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jp target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnp
Jumps to the label target if the parity flag is clear (PF = 0).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnp target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jz
je, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jz target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnz
jne, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnz target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jc
jb, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jc target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnae
jb, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnae target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnc
jae, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnc target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnb
jae, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnb target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnbe
ja, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnbe target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jna
jbe, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jna target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnge
jl, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnge target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnl
jge, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnl target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jng
jle, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jng target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnle
jg, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnle target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jpe
jp, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jpe target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jpo
jnp, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jpo target | target: int | emits Rel32Instr2 | std.x86_64.impl |
ret
Returns: pops an address and jumps to it.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
ret | emits Byte | std.x86_64.impl |
shl
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
shl rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = rd << imm. | std.x86_64.impl |
shl rd, imm | rd: Reg, imm: int | rd = rd << imm. | ||
shl rd, imm | rd: Reg32, imm: int | rd = rd << imm, in 32 bits. |
shl_cl
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
shl_cl(rd, w) | rd: Reg, w: int | emits Bytes<...> | rd = rd << cl. | std.x86_64.impl |
shl rd, cl | rd: Reg | rd = rd << cl. | ||
shl rd, cl | rd: Reg32 | rd = rd << cl, in 32 bits. |
shr
rd = rd >> imm, shifting in zeros.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
shr rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | std.x86_64.impl |
shr_cl
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
shr_cl(rd, w) | rd: Reg, w: int | emits Bytes<...> | rd = rd >> cl, shifting in zeros. | std.x86_64.impl |
shr rd, cl | rd: Reg | rd = rd >> cl, shifting in zeros. | ||
shr rd, cl | rd: Reg32 | rd = rd >> cl, shifting in zeros, in 32 bits. |
sar
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
sar rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = rd >> imm, shifting in copies of the sign bit. | std.x86_64.impl |
sar rd, imm | rd: Reg, imm: int | rd = rd >> imm, shifting in copies of the sign bit. | ||
sar rd, imm | rd: Reg32, imm: int | rd = rd >> imm, shifting in copies of the sign bit, in 32 bits. |
sar_cl
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
sar_cl(rd, w) | rd: Reg, w: int | emits Bytes<...> | rd = rd >> cl, shifting in copies of the sign bit. | std.x86_64.impl |
sar rd, cl | rd: Reg | rd = rd >> cl, shifting in copies of the sign bit. | ||
sar rd, cl | rd: Reg32 | rd = rd >> cl, shifting in copies of the sign bit, in 32 bits. |
lea
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
lea rd, src, w | rd: Reg, src: MemOperand, w: int | emits Bytes<...> | Loads the address src stands for into rd, without reading memory. | std.x86_64.impl |
lea rd, src, w | rd: Reg, src: RipLabel, w: int | Loads the address of the label src into rd. | std.x86_64.impl |
rip_label_instr
An instruction with opcode opcode whose memory operand is the label
addr, and whose other operand is r. The dialects’ [rel label]
forms are built on it.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rip_label_instr opcode, addr, r, w | opcode: int, addr: RipLabel, r: Reg, w: int | std.x86_64.impl |
push
Pushes the 64-bit rd onto the stack.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
push rd | rd: Reg | emits Bytes<...> | std.x86_64.impl |
pop
Pops 64 bits off the stack into rd.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
pop rd | rd: Reg | emits Bytes<...> | std.x86_64.impl |
syscall
Calls the operating system. On Linux, rax holds the call number and
rdi, rsi, rdx, … its arguments.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
syscall | emits Bytes<2> | std.x86_64.impl |
assert_valid_reg
Fails to compile unless r is a register number, 0 to 15. The
instructions here use it to check their operands.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
assert_valid_reg r | r: Reg |
mov_load_base
| Syntax | Parameters | Result | Description |
|---|---|---|---|
mov rd, [base] | rd: Reg, base: Reg | Loads the value at [base] into rd. | |
mov rd, [base] | rd: Reg32, base: Reg | Loads the 32-bit value at [base] into rd. |
mov_load_base_disp
| Syntax | Parameters | Result | Description |
|---|---|---|---|
mov rd, [base+disp] | rd: Reg, base: Reg, disp: int | Loads the value at [base + disp] into rd. | |
mov rd, [base+disp] | rd: Reg32, base: Reg, disp: int | Loads the 32-bit value at [base + disp] into rd. |
mov_load_indexed
| Syntax | Parameters | Result | Description |
|---|---|---|---|
mov rd, [base+index*scale+disp] | rd: Reg, base: Reg, index: Reg, scale: int, disp: int | Loads the value at [base + index * scale + disp] into rd. | |
mov rd, [base+index*scale+disp] | rd: Reg32, base: Reg, index: Reg, scale: int, disp: int | Loads the 32-bit value at [base + index * scale + disp] into rd. |
mov_load_rip
| Syntax | Parameters | Result | Description |
|---|---|---|---|
mov rd, [rip+disp] | rd: Reg, disp: int | Loads the value at [rip + disp], disp bytes past the end of the instruction, into rd. | |
mov rd, [rip+disp] | rd: Reg32, disp: int | Loads the 32-bit value at [rip + disp], disp bytes past the end of the instruction, into rd. |
mov_store_base
| Syntax | Parameters | Result | Description |
|---|---|---|---|
mov [base], rs | base: Reg, rs: Reg | Stores rs at [base]. | |
mov [base], rs | base: Reg, rs: Reg32 | Stores the 32-bit rs at [base]. |
mov_store_base_disp
| Syntax | Parameters | Result | Description |
|---|---|---|---|
mov [base+disp], rs | base: Reg, disp: int, rs: Reg | Stores rs at [base + disp]. | |
mov [base+disp], rs | base: Reg, disp: int, rs: Reg32 | Stores the 32-bit rs at [base + disp]. |
mov_store_indexed
| Syntax | Parameters | Result | Description |
|---|---|---|---|
mov [base+index*scale+disp], rs | base: Reg, index: Reg, scale: int, disp: int, rs: Reg | Stores rs at [base + index * scale + disp]. | |
mov [base+index*scale+disp], rs | base: Reg, index: Reg, scale: int, disp: int, rs: Reg32 | Stores the 32-bit rs at [base + index * scale + disp]. |
mov_store_rip
| Syntax | Parameters | Result | Description |
|---|---|---|---|
mov [rip+disp], rs | disp: int, rs: Reg | Stores rs at [rip + disp], disp bytes past the end of the instruction. | |
mov [rip+disp], rs | disp: int, rs: Reg32 | Stores the 32-bit rs at [rip + disp], disp bytes past the end of the instruction. |
lea_base
| Syntax | Parameters | Result | Description |
|---|---|---|---|
lea rd, [base] | rd: Reg, base: Reg | Loads the address [base] into rd, without reading memory. | |
lea rd, [base] | rd: Reg32, base: Reg | Loads the address [base] into the 32-bit rd, without reading memory. |
lea_base_disp
| Syntax | Parameters | Result | Description |
|---|---|---|---|
lea rd, [base+disp] | rd: Reg, base: Reg, disp: int | Loads the address [base + disp] into rd, without reading memory. | |
lea rd, [base+disp] | rd: Reg32, base: Reg, disp: int | Loads the address [base + disp] into the 32-bit rd, without reading memory. |
lea_indexed
| Syntax | Parameters | Result | Description |
|---|---|---|---|
lea rd, [base+index*scale+disp] | rd: Reg, base: Reg, index: Reg, scale: int, disp: int | Loads the address [base + index * scale + disp] into rd, without reading memory. | |
lea rd, [base+index*scale+disp] | rd: Reg32, base: Reg, index: Reg, scale: int, disp: int | Loads the address [base + index * scale + disp] into the 32-bit rd, without reading memory. |
lea_rip
| Syntax | Parameters | Result | Description |
|---|---|---|---|
lea rd, [rip+disp] | rd: Reg, disp: int | Loads the address [rip + disp], disp bytes past the end of the instruction, into rd. | |
lea rd, [rip+disp] | rd: Reg32, disp: int | Loads the address [rip + disp], disp bytes past the end of the instruction, into the 32-bit rd. |
shr_imm
| Syntax | Parameters | Result | Description |
|---|---|---|---|
shr rd, imm | rd: Reg, imm: int | rd = rd >> imm, shifting in zeros. | |
shr rd, imm | rd: Reg32, imm: int | rd = rd >> imm, shifting in zeros, in 32 bits. |
Re-exported
From std.x86_64.impl: Reg, r0, r1, r2, r3, r4, r5, r6, r7, r8, r9, r10, r11, r12, r13, r14, r15, rax, rcx, rdx, rbx, rsp, rbp, rsi, rdi, Reg32, eax, ecx, edx, ebx, esp, ebp, esi, edi, r8d, r9d, r10d, r11d, r12d, r13d, r14d, r15d, Byte, Bytes, MemBase, MemSib, MemRip, MemOperand, Rel32Instr, Rel32Instr2, RipLabel, RipRelInstr, RipRelInstrRex.
nasm
NASM-flavored x86-64 assembly: everything std.x86_64.intel provides,
plus the NASM spellings it doesn’t have.
from std.x86_64.nasm import *
lea rsi, [rel msg]
mov eax, [rel counter]
add [rel total], rcx
msg:
db "Hello, World!", 10
counter:
dd 0
total:
dq 0
[rip+disp] from std.x86_64.intel takes a displacement you already
know; [rel label] takes a label and computes the displacement once the
code is laid out. The two can’t share a spelling: a label is a plain int
too, so [rip+label] would be exactly as well-typed as [rip+disp].
Not covered: NASM’s $/$$, equ, times, resb, and
global/extern (use const, @for, and pub labels instead).
Re-exports std.x86_64.intel.
Macros
reg_field
The low 3 bits of r’s number: the part a ModRM or SIB field, or an
opcode, holds.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
reg_field(r) | r: Reg | returns int | std.x86_64.impl |
reg_ext
Bit 3 of r’s number, which goes in the REX prefix: 1 for r8 to
r15.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
reg_ext(r) | r: Reg | returns int | std.x86_64.impl |
rex_byte
A REX prefix, 0100WRXB. w selects 64-bit operands; r, x and b
are bit 3 of the ModRM.reg, SIB.index and ModRM.rm (or SIB.base, or
opcode) register numbers.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rex_byte(w, r, x, b) | w: int, r: int, x: int, b: int | returns Byte | std.x86_64.impl |
modrm_byte
A ModRM byte: mod (2 bits), reg (3) and rm (3).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
modrm_byte(mod, reg, rm) | mod: int, reg: int, rm: int | returns Byte | std.x86_64.impl |
sib_byte
A SIB byte: scale (2 bits, log2 of the index’s multiplier), index
(3) and base (3).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
sib_byte(scale, index, base) | scale: int, index: int, base: int | returns Byte | std.x86_64.impl |
byte_of
Byte index of value, counting from the least significant, byte 0.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
byte_of(value, index) | value: int, index: int | returns Byte | std.x86_64.impl |
mov
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
mov rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | rd = rs. | std.x86_64.impl |
mov rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = imm. With w = 1 it takes a full 64-bit imm. | std.x86_64.impl |
mov dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | Stores rs at dst. | std.x86_64.impl |
mov rd, src, w | rd: Reg, src: MemOperand, w: int | emits Bytes<...> | Loads the value at src into rd. | std.x86_64.impl |
mov dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | Stores imm at dst: 32 bits, sign-extended to 64 when w is 1. | std.x86_64.impl |
mov rd, src, w | rd: Reg, src: RipLabel, w: int | Loads the value at the label src into rd. | std.x86_64.impl | |
mov dst, rs, w | dst: RipLabel, rs: Reg, w: int | Stores rs at the label dst. | std.x86_64.impl | |
mov rd, rs | rd: Reg, rs: Reg | rd = rs. | std.x86_64.intel | |
mov rd, imm | rd: Reg, imm: int | rd = imm. | std.x86_64.intel | |
mov rd, rs | rd: Reg32, rs: Reg32 | rd = rs, in 32 bits. | std.x86_64.intel | |
mov rd, imm | rd: Reg32, imm: int | rd = imm, in 32 bits. | std.x86_64.intel | |
mov rd, src | rd: Reg, src: RipLabel | Loads the value at src into rd. | ||
mov rd, src | rd: Reg32, src: RipLabel | |||
mov dst, rs | dst: RipLabel, rs: Reg | Stores rs at dst. | ||
mov dst, rs | dst: RipLabel, rs: Reg32 |
add
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
add rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | rd = rd + rs. | std.x86_64.impl |
add rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = rd + imm, with a 32-bit imm sign-extended to 64 bits. | std.x86_64.impl |
add dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | [dst] = [dst] + rs. | std.x86_64.impl |
add dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | [dst] = [dst] + imm, with a 32-bit imm. | std.x86_64.impl |
add dst, rs, w | dst: RipLabel, rs: Reg, w: int | [dst] = [dst] + rs, where dst is a label. | std.x86_64.impl | |
add rd, rs | rd: Reg, rs: Reg | rd = rd + rs. | std.x86_64.intel | |
add rd, imm | rd: Reg, imm: int | rd = rd + imm. | std.x86_64.intel | |
add rd, rs | rd: Reg32, rs: Reg32 | rd = rd + rs, in 32 bits. | std.x86_64.intel | |
add rd, imm | rd: Reg32, imm: int | rd = rd + imm, in 32 bits. | std.x86_64.intel | |
add dst, rs | dst: RipLabel, rs: Reg | Adds rs to the value at dst. | ||
add dst, rs | dst: RipLabel, rs: Reg32 |
or
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
or rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | rd = rd | rs. | std.x86_64.impl |
or rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = rd | imm, with a 32-bit imm sign-extended to 64 bits. | std.x86_64.impl |
or dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | [dst] = [dst] | rs. | std.x86_64.impl |
or dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | [dst] = [dst] | imm, with a 32-bit imm. | std.x86_64.impl |
or dst, rs, w | dst: RipLabel, rs: Reg, w: int | [dst] = [dst] | rs, where dst is a label. | std.x86_64.impl | |
or rd, rs | rd: Reg, rs: Reg | rd = rd | rs. | std.x86_64.intel | |
or rd, imm | rd: Reg, imm: int | rd = rd | imm. | std.x86_64.intel | |
or rd, rs | rd: Reg32, rs: Reg32 | rd = rd | rs, in 32 bits. | std.x86_64.intel | |
or rd, imm | rd: Reg32, imm: int | rd = rd | imm, in 32 bits. | std.x86_64.intel | |
or dst, rs | dst: RipLabel, rs: Reg | ORs rs into the value at dst. | ||
or dst, rs | dst: RipLabel, rs: Reg32 |
and
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
and rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | rd = rd & rs. | std.x86_64.impl |
and rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = rd & imm, with a 32-bit imm sign-extended to 64 bits. | std.x86_64.impl |
and dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | [dst] = [dst] & rs. | std.x86_64.impl |
and dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | [dst] = [dst] & imm, with a 32-bit imm. | std.x86_64.impl |
and dst, rs, w | dst: RipLabel, rs: Reg, w: int | [dst] = [dst] & rs, where dst is a label. | std.x86_64.impl | |
and rd, rs | rd: Reg, rs: Reg | rd = rd & rs. | std.x86_64.intel | |
and rd, imm | rd: Reg, imm: int | rd = rd & imm. | std.x86_64.intel | |
and rd, rs | rd: Reg32, rs: Reg32 | rd = rd & rs, in 32 bits. | std.x86_64.intel | |
and rd, imm | rd: Reg32, imm: int | rd = rd & imm, in 32 bits. | std.x86_64.intel | |
and dst, rs | dst: RipLabel, rs: Reg | ANDs rs into the value at dst. | ||
and dst, rs | dst: RipLabel, rs: Reg32 |
sub
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
sub rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | rd = rd - rs. | std.x86_64.impl |
sub rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = rd - imm, with a 32-bit imm sign-extended to 64 bits. | std.x86_64.impl |
sub dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | [dst] = [dst] - rs. | std.x86_64.impl |
sub dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | [dst] = [dst] - imm, with a 32-bit imm. | std.x86_64.impl |
sub dst, rs, w | dst: RipLabel, rs: Reg, w: int | [dst] = [dst] - rs, where dst is a label. | std.x86_64.impl | |
sub rd, rs | rd: Reg, rs: Reg | rd = rd - rs. | std.x86_64.intel | |
sub rd, imm | rd: Reg, imm: int | rd = rd - imm. | std.x86_64.intel | |
sub rd, rs | rd: Reg32, rs: Reg32 | rd = rd - rs, in 32 bits. | std.x86_64.intel | |
sub rd, imm | rd: Reg32, imm: int | rd = rd - imm, in 32 bits. | std.x86_64.intel | |
sub dst, rs | dst: RipLabel, rs: Reg | Subtracts rs from the value at dst. | ||
sub dst, rs | dst: RipLabel, rs: Reg32 |
xor
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
xor rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | rd = rd ^ rs. | std.x86_64.impl |
xor rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = rd ^ imm, with a 32-bit imm sign-extended to 64 bits. | std.x86_64.impl |
xor dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | [dst] = [dst] ^ rs. | std.x86_64.impl |
xor dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | [dst] = [dst] ^ imm, with a 32-bit imm. | std.x86_64.impl |
xor dst, rs, w | dst: RipLabel, rs: Reg, w: int | [dst] = [dst] ^ rs, where dst is a label. | std.x86_64.impl | |
xor rd, rs | rd: Reg, rs: Reg | rd = rd ^ rs. | std.x86_64.intel | |
xor rd, imm | rd: Reg, imm: int | rd = rd ^ imm. | std.x86_64.intel | |
xor rd, rs | rd: Reg32, rs: Reg32 | rd = rd ^ rs, in 32 bits. | std.x86_64.intel | |
xor rd, imm | rd: Reg32, imm: int | rd = rd ^ imm, in 32 bits. | std.x86_64.intel | |
xor dst, rs | dst: RipLabel, rs: Reg | XORs rs into the value at dst. | ||
xor dst, rs | dst: RipLabel, rs: Reg32 |
cmp
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
cmp rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | Sets the flags from rd - rs, without storing it. | std.x86_64.impl |
cmp rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | Sets the flags from rd - imm, without storing it. | std.x86_64.impl |
cmp dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | Sets the flags from [dst] - rs, without storing it. | std.x86_64.impl |
cmp dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | Sets the flags from [dst] - imm, without storing it. | std.x86_64.impl |
cmp dst, rs, w | dst: RipLabel, rs: Reg, w: int | Sets the flags from [dst] - rs, where dst is a label, without storing it. | std.x86_64.impl | |
cmp rd, rs | rd: Reg, rs: Reg | Sets the flags from rd - rs, without storing it. | std.x86_64.intel | |
cmp rd, imm | rd: Reg, imm: int | Sets the flags from rd - imm, without storing it. | std.x86_64.intel | |
cmp rd, rs | rd: Reg32, rs: Reg32 | Sets the flags from rd - rs, in 32 bits, without storing it. | std.x86_64.intel | |
cmp rd, imm | rd: Reg32, imm: int | Sets the flags from rd - imm, in 32 bits, without storing it. | std.x86_64.intel | |
cmp dst, rs | dst: RipLabel, rs: Reg | Compares the value at dst with rs, setting flags like sub without storing. | ||
cmp dst, rs | dst: RipLabel, rs: Reg32 |
test
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
test rd, rs, w | rd: Reg, rs: Reg, w: int | emits Bytes<...> | Sets the flags from rd & rs, without storing it. | std.x86_64.impl |
test rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | Sets the flags from rd & imm, without storing it. | std.x86_64.impl |
test dst, rs, w | dst: MemOperand, rs: Reg, w: int | emits Bytes<...> | Sets the flags from [dst] & rs, without storing it. | std.x86_64.impl |
test dst, imm, w | dst: MemOperand, imm: int, w: int | emits Bytes<...> | Sets the flags from [dst] & imm, without storing it. | std.x86_64.impl |
test dst, rs, w | dst: RipLabel, rs: Reg, w: int | Sets the flags from [dst] & rs, where dst is a label, without storing it. | std.x86_64.impl | |
test rd, rs | rd: Reg, rs: Reg | Sets the flags from rd & rs, without storing it. | std.x86_64.intel | |
test rd, imm | rd: Reg, imm: int | Sets the flags from rd & imm, without storing it. | std.x86_64.intel | |
test rd, rs | rd: Reg32, rs: Reg32 | Sets the flags from rd & rs, in 32 bits, without storing it. | std.x86_64.intel | |
test rd, imm | rd: Reg32, imm: int | Sets the flags from rd & imm, in 32 bits, without storing it. | std.x86_64.intel | |
test dst, rs | dst: RipLabel, rs: Reg | ANDs the value at dst with rs, setting flags without storing. | ||
test dst, rs | dst: RipLabel, rs: Reg32 |
Mem
The memory at [base + disp]. A displacement from -128 to 127 takes one
byte; any other takes four.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
Mem(base, disp) | base: Reg, disp: int | returns MemOperand | std.x86_64.impl |
MemIndexed
The memory at [base + index * scale + disp]. scale is 1, 2, 4 or 8,
and index can be any register but rsp.
from std.x86_64.impl import *
mov rax, MemIndexed(rbx, r12, 4, 0), 1
4a 8b 04 a3
from std.x86_64.impl import *
mov rax, MemIndexed(rbx, rsp, 4, 0), 1
rsp cannot be a SIB index register
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
MemIndexed(base, index, scale, disp) | base: Reg, index: Reg, scale: int, disp: int | returns MemOperand | std.x86_64.impl |
MemRipRelative
The memory at [rip + disp]: disp bytes past the end of the
instruction. To address a label, use a RipLabel instead.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
MemRipRelative(disp) | disp: int | returns MemOperand | std.x86_64.impl |
jmp
Jumps to the label target.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jmp target | target: int | emits Rel32Instr | std.x86_64.impl |
call
Pushes the address of the next instruction and jumps to the label
target.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
call target | target: int | emits Rel32Instr | std.x86_64.impl |
je
Jumps to the label target if equal (ZF = 1).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
je target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jne
Jumps to the label target if not equal (ZF = 0).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jne target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jb
Jumps to the label target if below, unsigned (CF = 1).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jb target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jae
Jumps to the label target if above or equal, unsigned (CF = 0).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jae target | target: int | emits Rel32Instr2 | std.x86_64.impl |
ja
Jumps to the label target if above, unsigned.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
ja target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jbe
Jumps to the label target if below or equal, unsigned.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jbe target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jl
Jumps to the label target if less, signed.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jl target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jge
Jumps to the label target if greater or equal, signed.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jge target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jle
Jumps to the label target if less or equal, signed.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jle target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jg
Jumps to the label target if greater, signed.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jg target | target: int | emits Rel32Instr2 | std.x86_64.impl |
js
Jumps to the label target if the result was negative (SF = 1).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
js target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jns
Jumps to the label target if the result wasn’t negative (SF = 0).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jns target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jo
Jumps to the label target on signed overflow (OF = 1).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jo target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jno
Jumps to the label target without signed overflow (OF = 0).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jno target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jp
Jumps to the label target if the parity flag is set (PF = 1).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jp target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnp
Jumps to the label target if the parity flag is clear (PF = 0).
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnp target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jz
je, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jz target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnz
jne, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnz target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jc
jb, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jc target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnae
jb, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnae target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnc
jae, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnc target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnb
jae, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnb target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnbe
ja, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnbe target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jna
jbe, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jna target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnge
jl, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnge target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnl
jge, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnl target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jng
jle, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jng target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jnle
jg, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jnle target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jpe
jp, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jpe target | target: int | emits Rel32Instr2 | std.x86_64.impl |
jpo
jnp, under another name.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
jpo target | target: int | emits Rel32Instr2 | std.x86_64.impl |
ret
Returns: pops an address and jumps to it.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
ret | emits Byte | std.x86_64.impl |
shl
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
shl rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = rd << imm. | std.x86_64.impl |
shl rd, imm | rd: Reg, imm: int | rd = rd << imm. | std.x86_64.intel | |
shl rd, imm | rd: Reg32, imm: int | rd = rd << imm, in 32 bits. | std.x86_64.intel |
shl_cl
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
shl_cl(rd, w) | rd: Reg, w: int | emits Bytes<...> | rd = rd << cl. | std.x86_64.impl |
shl rd, cl | rd: Reg | rd = rd << cl. | std.x86_64.intel | |
shl rd, cl | rd: Reg32 | rd = rd << cl, in 32 bits. | std.x86_64.intel |
shr
rd = rd >> imm, shifting in zeros.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
shr rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | std.x86_64.impl |
shr_cl
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
shr_cl(rd, w) | rd: Reg, w: int | emits Bytes<...> | rd = rd >> cl, shifting in zeros. | std.x86_64.impl |
shr rd, cl | rd: Reg | rd = rd >> cl, shifting in zeros. | std.x86_64.intel | |
shr rd, cl | rd: Reg32 | rd = rd >> cl, shifting in zeros, in 32 bits. | std.x86_64.intel |
sar
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
sar rd, imm, w | rd: Reg, imm: int, w: int | emits Bytes<...> | rd = rd >> imm, shifting in copies of the sign bit. | std.x86_64.impl |
sar rd, imm | rd: Reg, imm: int | rd = rd >> imm, shifting in copies of the sign bit. | std.x86_64.intel | |
sar rd, imm | rd: Reg32, imm: int | rd = rd >> imm, shifting in copies of the sign bit, in 32 bits. | std.x86_64.intel |
sar_cl
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
sar_cl(rd, w) | rd: Reg, w: int | emits Bytes<...> | rd = rd >> cl, shifting in copies of the sign bit. | std.x86_64.impl |
sar rd, cl | rd: Reg | rd = rd >> cl, shifting in copies of the sign bit. | std.x86_64.intel | |
sar rd, cl | rd: Reg32 | rd = rd >> cl, shifting in copies of the sign bit, in 32 bits. | std.x86_64.intel |
lea
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
lea rd, src, w | rd: Reg, src: MemOperand, w: int | emits Bytes<...> | Loads the address src stands for into rd, without reading memory. | std.x86_64.impl |
lea rd, src, w | rd: Reg, src: RipLabel, w: int | Loads the address of the label src into rd. | std.x86_64.impl | |
lea rd, src | rd: Reg, src: RipLabel | Loads the address of src into rd. | ||
lea rd, src | rd: Reg32, src: RipLabel |
rip_label_instr
An instruction with opcode opcode whose memory operand is the label
addr, and whose other operand is r. The dialects’ [rel label]
forms are built on it.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
rip_label_instr opcode, addr, r, w | opcode: int, addr: RipLabel, r: Reg, w: int | std.x86_64.impl |
push
Pushes the 64-bit rd onto the stack.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
push rd | rd: Reg | emits Bytes<...> | std.x86_64.impl |
pop
Pops 64 bits off the stack into rd.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
pop rd | rd: Reg | emits Bytes<...> | std.x86_64.impl |
syscall
Calls the operating system. On Linux, rax holds the call number and
rdi, rsi, rdx, … its arguments.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
syscall | emits Bytes<2> | std.x86_64.impl |
assert_valid_reg
Fails to compile unless r is a register number, 0 to 15. The
instructions here use it to check their operands.
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
assert_valid_reg r | r: Reg | std.x86_64.intel |
mov_load_base
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
mov rd, [base] | rd: Reg, base: Reg | Loads the value at [base] into rd. | std.x86_64.intel | |
mov rd, [base] | rd: Reg32, base: Reg | Loads the 32-bit value at [base] into rd. | std.x86_64.intel |
mov_load_base_disp
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
mov rd, [base+disp] | rd: Reg, base: Reg, disp: int | Loads the value at [base + disp] into rd. | std.x86_64.intel | |
mov rd, [base+disp] | rd: Reg32, base: Reg, disp: int | Loads the 32-bit value at [base + disp] into rd. | std.x86_64.intel |
mov_load_indexed
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
mov rd, [base+index*scale+disp] | rd: Reg, base: Reg, index: Reg, scale: int, disp: int | Loads the value at [base + index * scale + disp] into rd. | std.x86_64.intel | |
mov rd, [base+index*scale+disp] | rd: Reg32, base: Reg, index: Reg, scale: int, disp: int | Loads the 32-bit value at [base + index * scale + disp] into rd. | std.x86_64.intel |
mov_load_rip
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
mov rd, [rip+disp] | rd: Reg, disp: int | Loads the value at [rip + disp], disp bytes past the end of the instruction, into rd. | std.x86_64.intel | |
mov rd, [rip+disp] | rd: Reg32, disp: int | Loads the 32-bit value at [rip + disp], disp bytes past the end of the instruction, into rd. | std.x86_64.intel |
mov_store_base
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
mov [base], rs | base: Reg, rs: Reg | Stores rs at [base]. | std.x86_64.intel | |
mov [base], rs | base: Reg, rs: Reg32 | Stores the 32-bit rs at [base]. | std.x86_64.intel |
mov_store_base_disp
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
mov [base+disp], rs | base: Reg, disp: int, rs: Reg | Stores rs at [base + disp]. | std.x86_64.intel | |
mov [base+disp], rs | base: Reg, disp: int, rs: Reg32 | Stores the 32-bit rs at [base + disp]. | std.x86_64.intel |
mov_store_indexed
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
mov [base+index*scale+disp], rs | base: Reg, index: Reg, scale: int, disp: int, rs: Reg | Stores rs at [base + index * scale + disp]. | std.x86_64.intel | |
mov [base+index*scale+disp], rs | base: Reg, index: Reg, scale: int, disp: int, rs: Reg32 | Stores the 32-bit rs at [base + index * scale + disp]. | std.x86_64.intel |
mov_store_rip
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
mov [rip+disp], rs | disp: int, rs: Reg | Stores rs at [rip + disp], disp bytes past the end of the instruction. | std.x86_64.intel | |
mov [rip+disp], rs | disp: int, rs: Reg32 | Stores the 32-bit rs at [rip + disp], disp bytes past the end of the instruction. | std.x86_64.intel |
lea_base
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
lea rd, [base] | rd: Reg, base: Reg | Loads the address [base] into rd, without reading memory. | std.x86_64.intel | |
lea rd, [base] | rd: Reg32, base: Reg | Loads the address [base] into the 32-bit rd, without reading memory. | std.x86_64.intel |
lea_base_disp
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
lea rd, [base+disp] | rd: Reg, base: Reg, disp: int | Loads the address [base + disp] into rd, without reading memory. | std.x86_64.intel | |
lea rd, [base+disp] | rd: Reg32, base: Reg, disp: int | Loads the address [base + disp] into the 32-bit rd, without reading memory. | std.x86_64.intel |
lea_indexed
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
lea rd, [base+index*scale+disp] | rd: Reg, base: Reg, index: Reg, scale: int, disp: int | Loads the address [base + index * scale + disp] into rd, without reading memory. | std.x86_64.intel | |
lea rd, [base+index*scale+disp] | rd: Reg32, base: Reg, index: Reg, scale: int, disp: int | Loads the address [base + index * scale + disp] into the 32-bit rd, without reading memory. | std.x86_64.intel |
lea_rip
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
lea rd, [rip+disp] | rd: Reg, disp: int | Loads the address [rip + disp], disp bytes past the end of the instruction, into rd. | std.x86_64.intel | |
lea rd, [rip+disp] | rd: Reg32, disp: int | Loads the address [rip + disp], disp bytes past the end of the instruction, into the 32-bit rd. | std.x86_64.intel |
shr_imm
| Syntax | Parameters | Result | Description | From |
|---|---|---|---|---|
shr rd, imm | rd: Reg, imm: int | rd = rd >> imm, shifting in zeros. | std.x86_64.intel | |
shr rd, imm | rd: Reg32, imm: int | rd = rd >> imm, shifting in zeros, in 32 bits. | std.x86_64.intel |
rel
[rel label]: the memory at label, addressed relative to the next
instruction, so the code works wherever it’s loaded.
from std.x86_64.nasm import *
lea rsi, [rel msg] # msg is 6 bytes past the end of this instruction
mov eax, [rel msg] # and right after this one
msg:
db "Hi", 10
48 8d 35 06 00 00 00
8b 05 00 00 00 00
48 69 0a
| Syntax | Parameters | Result | Description |
|---|---|---|---|
[rel target] | target: int | returns RipLabel |
data_int
value as width little-endian bytes, truncated the way NASM does.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
data_int(value, width) | value: int, width: int | returns Bytes<...> |
data_string
A string literal’s UTF-8 bytes, with no terminator.
| Syntax | Parameters | Result | Description |
|---|---|---|---|
data_string(source) | <S>, source: S | returns Bytes<...> |
db
| Syntax | Parameters | Result | Description |
|---|---|---|---|
db a | a: int | Bytes. Each operand is an integer, stored in one byte, or a string literal, stored as its UTF-8 bytes with no terminator (add , 0 yourself). Up to four operands; db "Hello", 10 is the common case. | |
db a | <S>, a: S | ||
db a, b | <S>, a: S, b: int | ||
db a, b, c | <S>, a: S, b: int, c: int | ||
db a, b | a: int, b: int | ||
db a, b, c | a: int, b: int, c: int | ||
db a, b, c, d | a: int, b: int, c: int, d: int |
dw
| Syntax | Parameters | Result | Description |
|---|---|---|---|
dw a | a: int | 16-bit little-endian words, one or two per line. | |
dw a, b | a: int, b: int |
dd
| Syntax | Parameters | Result | Description |
|---|---|---|---|
dd a | a: int | 32-bit little-endian doublewords, one or two per line. | |
dd a, b | a: int, b: int |
dq
| Syntax | Parameters | Result | Description |
|---|---|---|---|
dq a | a: int | 64-bit little-endian quadwords, one or two per line. | |
dq a, b | a: int, b: int |
Re-exported
From std.x86_64.impl: Reg, r0, r1, r2, r3, r4, r5, r6, r7, r8, r9, r10, r11, r12, r13, r14, r15, rax, rcx, rdx, rbx, rsp, rbp, rsi, rdi, Reg32, eax, ecx, edx, ebx, esp, ebp, esi, edi, r8d, r9d, r10d, r11d, r12d, r13d, r14d, r15d, Byte, Bytes, MemBase, MemSib, MemRip, MemOperand, Rel32Instr, Rel32Instr2, RipLabel, RipRelInstr, RipRelInstrRex.
Writing the docs
This book lives next to the compiler, so a change to the language and its
documentation land in the same pull request. Its examples are tests:
cargo test --test book compiles every one of them, so an example that
stops working fails CI instead of quietly going stale.
How pages are organized
Each chapter starts with a general page: what the feature is for, and a short example. Its subchapters go deep on one piece each. A subchapter usually has:
- A one-line summary.
- A Syntax block, marked
ignore. - A small, complete, tested example with its output.
- Details, edge cases and errors, each with its own example where possible.
Code blocks
| Fence | Meaning |
|---|---|
```basm | Must compile, with every lint except generated_declarations denied. |
```basm,fail | Must fail to compile. |
```basm,ignore | A fragment: highlighted, but not compiled. |
```basm,file=name.basm | Not compiled itself. Saved as name.basm next to the page’s later examples, so they can import it with from .name import .... |
Directly after a compiled basm block, any of these check its result:
| Fence | Checks |
|---|---|
```emits | The emitted values, in order. |
```bytes | The bytes bitter encode packs them into, in hex. |
```error | Text the compile error must contain (after basm,fail). |
In an emits block, integers can share a line, separated by spaces. Other
values go one per line, written as the test prints them:
Point { x: 1, y: 2 }, bits<8> { value: 65 }, Shape.Circle(2),
Word<16, 1> { ... } (an enum argument is shown as its index), or
<name> for a label imported from another file. If you’re unsure, write
your best guess: the test failure shows the actual output.
Writing examples
Each basm block is compiled on its own, from the repository root, so it can
import std, and any file= blocks earlier on the same page.
Top-level @emit isn’t allowed, so an example that shows values defines a
small macro to emit them:
macro show(value: int) {
@emit value
}
show 6 * 7
42
Prefer a complete, checked example over an ignored fragment. Keep ignore
for syntax summaries, and for code that can’t stand alone.
When the language has a limitation or a known bug, say so in a note, and
show it with a fail example if you can. When the bug is fixed, the test
fails, which reminds you to update the page.
Building
mdbook serve docs/book --open # live preview
mdbook build docs/book # writes docs/book/book
The introduction is included from the README’s overview anchor, so edit it
there.
The std reference
std/reference/ is generated from std’s doc comments, so don’t edit it by
hand. Edit the ## and #! comments in std/, then regenerate it:
bitterasm doc std -o docs/book/src/std/reference --summary docs/book/src/SUMMARY.md
That also updates the reference’s entries in SUMMARY.md, so a new module
needs nothing else. cargo test --test doc_comments fails while either is
out of date.