Write a small tokenizer and explain the design choices.
A lexer converts a character stream into tokens. A hand-written tokenizer scans one character at a time and skips whitespace.
def tokenize(src):
tokens, i = [], 0
while i < len(src):
c = src[i]
if c.isspace():
i += 1
elif c.isdigit():
j = i
while j < len(src) and src[j].isdigit():
j += 1
tokens.append(("NUM", int(src[i:j])))
i = j
elif c.isalpha():
j = i
while j < len(src) and src[j].isalnum():
j += 1
tokens.append(("IDENT", src[i:j]))
i = j
elif c in "+-*/()":
tokens.append((c, c))
i += 1
else:
raise SyntaxError(f"unexpected {c!r} at {i}")
tokens.append(("EOF", None))
return tokens
Real lexers use maximal munch, precedence for multi-character operators such as == over =, and a DFA or a generator like flex. Track line and column numbers so later phases can report precise errors.