@@ -55,11 +55,10 @@ The pipeline for a single file is: `String` → **parse** → `AST` → **check*
5555has its own README worth reading before making non-trivial changes there (` src/parse/README.md ` ,
5656` src/check/README.md ` , ` src/generate/README.md ` ).
5757
58- - ** ` src/parse ` ** : ` lex/ ` tokenizes source into a ` Vec<Token> ` , handling indentation by emitting synthetic
59- ` Indent ` /` Dedent ` tokens that the parser uses to delimit blocks. The parser (` expression.rs ` , ` statement.rs ` ,
58+ - ** ` src/parse ` ** : ` lex/ ` tokenizes source into a ` Vec<Token> ` . The parser (` expression.rs ` , ` statement.rs ` ,
6059 ` class.rs ` , ` definition.rs ` , ` control_flow_*.rs ` , ` collection.rs ` , ` call.rs ` , ` operation.rs ` , ` ty.rs ` , etc.)
6160 walks tokens via a ` TokenIterator ` and builds an ` AST ` (` parse/ast ` ), where each node carries a ` Position ` for
62- error reporting. Errors here mean the token stream doesn't conform to the grammar.
61+ error reporting.
6362
6463- ** ` src/check ` ** : the type checker, and where most language semantics live. Three phases:
6564 1 . ** Context building** (` check/context ` ): scans all ASTs (including implicit/explicit imports and built-in
@@ -102,6 +101,54 @@ has its own README worth reading before making non-trivial changes there (`src/p
102101 fixture path resolution, randomized temp output dirs, and the Python-AST-diff assertion logic used by
103102 ` test_directory ` /` test_directory_args ` .
104103
104+ ## Block syntax (post indent/dedent removal)
105+
106+ ` for ` /` while ` /` with ` bodies always require an explicit ` do ... end ` block — there is no single-statement
107+ shorthand for these three (` for a in b do c ` is a parse error; it must be ` for a in b do c end ` ). ` if ` /` then ` /
108+ ` else ` branches are the exception: each branch is parsed as one ` parse_expr_or_stmt ` , which accepts either a
109+ bare single statement/expression or an explicit ` do ... end ` block (` if a then do ... end else c ` is valid).
110+ A leading newline before a statement/expression is insignificant whitespace and is skipped (see
111+ ` parse_expression ` 's and ` parse_expr_or_stmt ` 's ` eat_while(&Token::NL) ` ) — but a * trailing* newline is still
112+ usually required as a statement separator inside a block, so constructs that need to look past it for an
113+ optional following keyword (e.g. ` parse_if ` scanning past newlines for a possible ` else ` ) must use a
114+ lookahead-with-rollback helper (` LexIterator::peek_if_skipping ` ) rather than unconditionally consuming the
115+ newline, or they'll break "no ` else ` , followed by more statements in the same block".
116+
117+ The call-site "handle" construct for a call that may raise is ` <expr> ! where <case> ... end ` , e.g.
118+ ` f(10) ! where err: MyErr => do ... end end ` — the ` ! ` marks the call as fallible and must be consumed before
119+ looking for ` where ` (` parse_expr_or_stmt ` in ` expr_or_stmt.rs ` ).
120+
121+ ` type X: Parent when <cond> ` (single-line) / ` type X: Parent when\n <cond>\n...\nend ` (multi-line, terminated
122+ by ` end ` ) is the * conditional type alias* form (produces ` Node::TypeAlias ` , binds ` self ` to ` Parent ` while
123+ checking the conditions). ` type X where <defs> end ` is a different form — an interface/type body of field and
124+ function * signatures* (produces ` Node::TypeDef ` , does ** not** bind ` self ` ). These two are easy to conflate
125+ (` when ` vs ` where ` ) since both start with ` type X: Parent ` ; picking the wrong one either fails to parse or fails
126+ type-checking with a confusing "Undefined variable: self".
127+
128+ ## Known incomplete work (branch ` feat-remove-indent-dedent ` , as of 2026-08-24)
129+
130+ This branch is mid-refactor from indentation-based blocks to the ` do ` /` end ` scheme above, and several
131+ ` tests/resource/valid/** ` fixtures were rewritten ahead of the features they exercise:
132+
133+ - ** ` trait ` is unimplemented.** It's a real, documented keyword (see the README's "traits" section and
134+ ` docs/spec/trait-def ` in ` docs/spec/grammar.md ` ) but the lexer/parser has zero support for it today. A few
135+ fixtures (` tests/resource/valid/class/parent.mamba ` , ` multiple_parent.mamba ` ,
136+ ` fun_with_body_in_interface.mamba ` , ` class_super_one_line_init.mamba ` , and transitively ` types.mamba ` via a
137+ dropped parent class) were rewritten to use ` trait ` and no longer parse. Fixing these needs either
138+ implementing ` trait ` as a real parser+checker+codegen feature, or reverting them to the ` class ` /` type ` -based
139+ syntax their paired ` .py ` reference files still expect.
140+ - ** Class-body statements/field-initializers that depend on constructor state are never moved into a generated
141+ ` __init__ ` .** E.g. ` class X(a: Float) where\n def y: Y := Y(a)\nend ` (a bare, non-` def ` constructor arg used
142+ in a field initializer) or a bare executable statement in a class body (e.g. a ` print(...) ` call) — Python
143+ reference fixtures expect these to be hoisted into ` __init__ ` (with a ` None ` placeholder left at class level
144+ for fields), but ` src/generate/convert/class.rs ` 's ` extract_class ` /` init ` only handles parent-` __init__ `
145+ calls and auto-generated ` self.field = arg ` assignments for constructor args, not general relocation. This
146+ needs a free-variable analysis over ` Core ` expressions to detect which class-body statements reference
147+ constructor-only names. The type checker itself does correctly resolve these now (see ` constrain_class_body `
148+ in ` src/check/constrain/generate/class.rs ` , which binds non-` def ` constructor args into the class body's
149+ environment) — it's specifically the codegen relocation that's missing, so affected programs type-check but
150+ transpile to Python that references undefined names.
151+
105152## Documentation
106153
107154` docs/ ` contains the (partially outdated, per its own README) language specification and philosophy docs,
0 commit comments