Repository navigation
OSH Parser
andychu edited this page Nov 16, 2017
·
20 revisions
These facts are useful for the parsing contest.
- 15 lexer modes (lexical state)
- 233 IDs (token types / node types) in 23 kinds (
core/id_kind_test.pyshows this) - 3 recursive descent parsers (command, word,
[[) - 1 Pratt parser (arithmetic)
-
[fallback reusesosh/bool_parser - TODO: modify asdl.py to show these stats?
- X product types
- X sum types with X alternatives
- what CPython opcodes does it use?
- how many lines of code does it use in CPython? (Compare with execution.)
- What is the distribution of ASDL string and array lengths per node type?
- note: there are several uses of string, not just token. Is this a good or bad optimization?
- Brace detection -- this is a separate metaprogramming pass (doesn't depend on input). This is a recursive parser, although it operates entirely on token types and not chars/strings?
-
core/glob_.py-- escape and unescape - TODO: regex escape, for passing to regcomp()
-
core/word_eval.py-- after evaluating VarOp arguments, we compile globs to Python regexes, e.g. for${x%foo*} - IFS splitting (this is quite slow and needs to be sped up!)
-
core/args.py-- this is not a recursive parser -
echo -e-- backslash escapes (andprintfif it turns out we need it as a builtin) -
readwithout -r -- backslash escapes are parsed
- Polymorphism:
- FileLineReader, StringLineReader, VirtualLineReader for here docs.
-
BoolParsercan taketest_builtin._StringWordEmitterorWordParser