← Back to Home

Twitter Sentiment Analysis

Bilingual Deep-Learning NLP (Indonesian + English)

Bilingual Deep-Learning NLP (Indonesian + English)

#machine learning

// Overview

A pair of deep-learning sentiment classifiers for tweets — one for Indonesian, one for English — that clean raw, noisy tweet text, embed it with trained word vectors, and predict sentiment with a CNN + LSTM network, served behind a simple Flask web API. (Built during an internship at Kazee.)

Background

▼ more▲ less

Social-media text is a hard input for classic ML:

  • Tweets are noisy. Mentions, links, hashtags, emoji, slang, and repeated letters drown the actual signal and have to be stripped before modeling.

  • Language matters. Indonesian needs its own stemming/stopword handling (Sastrawi); reusing an English pipeline would mangle it — hence two dedicated pipelines.

  • Word order carries sentiment. Bag-of-words loses negation and phrasing, motivating sequence models (CNN for local n-gram features, LSTM for order) over plain counting.

  • Pretrained vectors help. Training Word2Vec on the tweet corpus gives denser, more meaningful inputs than one-hot encoding.

  • Needs to be callable. The classifier had to be usable as a service, not just a notebook.

Solution

▼ more▲ less
  • Tweet cleaning. Regex-based removal of RT markers, mentions, URLs, hashtags, numbers, repeated characters, and non-letter symbols, then lowercase + tokenize.

  • Language-specific normalization. Indonesian → Sastrawi stopword removal + stemming (with an extended stopword list); English → NLTK tokenization.

  • Sequence encoding. A fitted Keras Tokenizer maps text to integer sequences, padded to a fixed max length.

  • Trained embeddings. Word2Vec vectors trained on the corpus initialize the embedding layer (Doc2Vec explored as an alternative representation).

  • CNN + LSTM model. Stacked Conv1D blocks (with dropout) extract local features, an LSTM captures sequence/order, and Dense + softmax produces the 3-class sentiment output; trained with checkpointing on best validation accuracy.

  • Flask serving. A small Flask app exposes a /predict endpoint that runs the full clean→encode→predict pipeline on incoming text.

Impact

▼ more▲ less
  • End-to-end working classifiers for two languages, from raw tweet to served prediction.

  • Proper Indonesian NLP, not an English pipeline bolted onto Indonesian text.

  • Sequence-aware modeling (CNN+LSTM) that respects word order and local phrasing.

  • Reusable as a service via a simple HTTP endpoint.

  • A reproducible workflow — preprocessing, embedding training, model training, and reporting captured as notebooks.

// Tech Stack

LayerTechnologies
ModelingKeras / TensorFlow — Embedding → Conv1D (×3, dropout) → LSTM → Dense → softmax
EmbeddingsWord2Vec (Gensim) trained on corpus; Doc2Vec explored
Preprocessing (ID)Sastrawi stemmer + stopword remover, regex tweet cleaning
Preprocessing (EN)NLTK tokenization, regex tweet cleaning
EncodingKeras Tokenizer + pad_sequences
ServingFlask (/predict endpoint)
WorkflowJupyter notebooks (preprocess · embeddings · model · report), pickled tokenizer + saved model/weights

// Architecture & Diagrams

No architecture diagrams configured yet.

Disclaimer: sample visuals may contain anonymized, simulated, or non-production values for presentation purposes.

// Demo

No demo configured yet.