Bilingual Deep-Learning NLP (Indonesian + English)
Bilingual Deep-Learning NLP (Indonesian + English)
A pair of deep-learning sentiment classifiers for tweets — one for Indonesian, one for English — that clean raw, noisy tweet text, embed it with trained word vectors, and predict sentiment with a CNN + LSTM network, served behind a simple Flask web API. (Built during an internship at Kazee.)
Social-media text is a hard input for classic ML:
Tweets are noisy. Mentions, links, hashtags, emoji, slang, and repeated letters drown the actual signal and have to be stripped before modeling.
Language matters. Indonesian needs its own stemming/stopword handling (Sastrawi); reusing an English pipeline would mangle it — hence two dedicated pipelines.
Word order carries sentiment. Bag-of-words loses negation and phrasing, motivating sequence models (CNN for local n-gram features, LSTM for order) over plain counting.
Pretrained vectors help. Training Word2Vec on the tweet corpus gives denser, more meaningful inputs than one-hot encoding.
Needs to be callable. The classifier had to be usable as a service, not just a notebook.
Tweet cleaning. Regex-based removal of RT markers, mentions, URLs, hashtags, numbers, repeated characters, and non-letter symbols, then lowercase + tokenize.
Language-specific normalization. Indonesian → Sastrawi stopword removal + stemming (with an extended stopword list); English → NLTK tokenization.
Sequence encoding. A fitted Keras Tokenizer maps text to integer sequences, padded to a fixed max length.
Trained embeddings. Word2Vec vectors trained on the corpus initialize the embedding layer (Doc2Vec explored as an alternative representation).
CNN + LSTM model. Stacked Conv1D blocks (with dropout) extract local features, an LSTM captures sequence/order, and Dense + softmax produces the 3-class sentiment output; trained with checkpointing on best validation accuracy.
Flask serving. A small Flask app exposes a /predict endpoint that runs the full
clean→encode→predict pipeline on incoming text.
End-to-end working classifiers for two languages, from raw tweet to served prediction.
Proper Indonesian NLP, not an English pipeline bolted onto Indonesian text.
Sequence-aware modeling (CNN+LSTM) that respects word order and local phrasing.
Reusable as a service via a simple HTTP endpoint.
A reproducible workflow — preprocessing, embedding training, model training, and reporting captured as notebooks.
| Layer | Technologies |
|---|---|
| Modeling | Keras / TensorFlow — Embedding → Conv1D (×3, dropout) → LSTM → Dense → softmax |
| Embeddings | Word2Vec (Gensim) trained on corpus; Doc2Vec explored |
| Preprocessing (ID) | Sastrawi stemmer + stopword remover, regex tweet cleaning |
| Preprocessing (EN) | NLTK tokenization, regex tweet cleaning |
| Encoding | Keras Tokenizer + pad_sequences |
| Serving | Flask (/predict endpoint) |
| Workflow | Jupyter notebooks (preprocess · embeddings · model · report), pickled tokenizer + saved model/weights |
Disclaimer: sample visuals may contain anonymized, simulated, or non-production values for presentation purposes.