Character-Level Deep Learning
Character-Level Deep Learning
A deep-learning model that predicts a person's gender from their name alone, by reading the name character by character through a CNN + LSTM network — no name dictionary required — served behind a simple Flask web API. (Built during an internship at Kazee.)
Predicting gender from a name is deceptively tricky:
Dictionaries don't generalize. A lookup table fails on rare, novel, or compound names; the signal really lives in how a name is spelled.
Sub-word patterns matter. Endings and letter combinations carry most of the gender signal, which a character-level model can learn directly.
Order counts. The same letters in different orders mean different things, motivating a sequence model (CNN for local patterns + LSTM for order) over a bag-of-characters.
Short, variable input. Names are short and vary in length, so fixed-length padding and a compact network were appropriate.
Needs to be callable. The classifier had to run as a service, not just a notebook.
Character preprocessing. Lowercase the name and split it into a list of characters.
Sequence encoding. A fitted character tokenizer maps characters to integers; sequences are padded to a fixed max length (40).
Character embedding. A learned Embedding layer turns character ids into dense vectors.
CNN + LSTM model. A Conv1D layer extracts local letter-pattern features, an LSTM models character order, and Dense → sigmoid produces the binary gender probability; an LSTM-stacked variant was also trained for comparison.
Flask serving. A small Flask app exposes a /predict_gender endpoint that runs the full
preprocess→encode→predict pipeline on an incoming name.
Dictionary-free prediction that generalizes to unseen names from spelling alone.
Character-level modeling that learns the sub-word patterns carrying the gender signal.
Sequence-aware (CNN+LSTM) rather than treating a name as an unordered bag of letters.
Reusable as a service via a simple HTTP endpoint.
Reproducible workflow — model training and prediction captured as notebooks with saved artifacts.
| Layer | Technologies |
|---|---|
| Modeling | Keras / TensorFlow — Embedding → Conv1D → LSTM → Dense → sigmoid (LSTM-stacked variant also trained) |
| Encoding | Character-level Keras Tokenizer + pad_sequences (maxlen 40) |
| Serving | Flask (/predict_gender endpoint) |
| Workflow | Jupyter notebooks (model · predict), pickled tokenizer + saved model/weights |
Disclaimer: sample visuals may contain anonymized, simulated, or non-production values for presentation purposes.