Skip to content

Client work · Live · OSINT · Data

OSINT Monitoring

A Telegram OSINT and monitoring platform. Scrapes channels through ghost accounts, indexes and translates messages, and exposes everything through a dashboard with search, charts, and spreadsheet exports.

Vue 3DjangoCeleryTelethonOpenSearchTerraformAWS

The problem

Analysts tracking disinformation, extremist networks, or geopolitical narratives on Telegram had no systematic way to monitor channels at scale. Manual monitoring doesn't scale past a handful of channels, and Telegram's API restrictions make automated collection non-trivial. Existing tools were either expensive enterprise platforms or fragile scripts that broke on rate limits.

The challenge was building something that could reliably scrape hundreds of channels through pools of real user accounts, handle Telegram's aggressive rate limiting and ban detection, and present the data in a way that analysts could actually work with: searchable, filterable, exportable, and translated.

Ghost account scraping

Pools of real Telegram accounts (Telethon sessions) join channels and pull message history on a schedule. Handles FloodWait rate limits, private invite links, and account churn.

Full-text search

Every message is bulk-indexed into OpenSearch. Analysts can search across all scraped content with filters for channel, date range, tags, and engagement metrics.

Machine translation

Self-hosted LibreTranslate translates every exported message to English. Also exposed as an ad-hoc endpoint in the dashboard for quick lookups.

Dashboard and charts

Vue 3 SPA with per-channel drill-downs, activity charts (Chart.js), ghost user leaderboards, and multi-project tenancy with role-based access.

XLSX export pipeline

Async export jobs walk the filtered queryset, translate each message, and build formatted spreadsheets. Status machine tracks progress; the UI polls for completion.

Active measures

Beyond passive scraping, the platform can send messages into channels through ghost accounts, enabling two-way operational capability.

How it works

The platform runs on AWS (eu-central-1) with a Vue 3 SPA served from CloudFront, a Django REST API on ECS Fargate, Postgres on RDS, OpenSearch for full-text search, and RabbitMQ (Amazon MQ) connecting the scrape pipeline. All infrastructure is managed through modularized Terraform with remote state in S3.

01

Celery beat fires periodic scrape tasks per ghost-user / channel pair

02

Worker spins up Telethon client, joins channel, upserts messages to Postgres

03

Un-indexed messages are bulk-pushed to OpenSearch in batches of 500

04

Analysts query via the dashboard; full-text hits OpenSearch, filters hit Postgres

05

Export jobs translate each message via LibreTranslate and build XLSX files

The scrape loop is fully async: Celery beat fires periodic tasks per ghost-user / channel pair, workers spin up Telethon clients to pull messages, and results are upserted to Postgres with media going to S3. Un-indexed messages are bulk-pushed to OpenSearch in batches of 500. The export pipeline walks filtered querysets, translates each message via LibreTranslate, and builds XLSX files with openpyxl.

The stateless tiers (web and worker) autoscale on CPU and memory target-tracking policies. The SPA is served from CloudFront edge locations. A documented scale-to-zero tfvars configuration acts as a cost kill-switch.

Decisions & trade-offs

One container image, three roles

The same ECR image serves the web tier (gunicorn), the worker tier (celery worker), and the scheduler (celery beat), differentiated only by the container command. Slightly larger images, but drastically simpler CI/CD and rollbacks.

OpenSearch for text, Postgres for everything else

Full-text search is offloaded to OpenSearch so Postgres stays fast for structured queries. Only id, raw_text, and text are indexed, keeping the search cluster cheap. Richer analytics fall back to Postgres where the full data model lives.

Queue-based decoupling via RabbitMQ

Scraping is fully async. Web traffic and scrape load never block each other, and workers scale independently based on queue depth. This also isolates Telegram rate-limit backpressure from the user-facing API.

Heroku to AWS migration with Terraform

The entire AWS environment is modularized Terraform: VPC, ECS cluster, RDS, Amazon MQ, OpenSearch, S3/CloudFront, bastion host, and per-service ALBs. Remote state in S3, environment switching via tfvars, and a documented scale-to-zero kill switch for cost control.

Scaling realities

The hard ceiling isn't AWS. It's Telegram itself. Scraping happens through real user accounts via MTProto, and Telegram enforces per-account rate limits (FloodWait), join limits, and behavioral bans. Throughput scales with the number of healthy ghost accounts, not with worker count. Adding ECS tasks beyond the number of ghost users buys little.

Each scrape task spins up a fresh asyncio event loop and Telethon client per invocation. Fine at dozens of channels, wasteful at thousands. A persistent-client model with long-lived workers each owning a pool of sessions would be the next big lever. The export pipeline translates sequentially, one HTTP call per message, so large exports are slow. Batching or pre-translating at ingest time would scale far better.

My role

Everything. I designed the product, built the Vue 3 dashboard, wrote the Django backend and Celery pipeline, set up the OpenSearch indexing, integrated Telethon for the scraping layer, configured LibreTranslate for machine translation, and managed the full Heroku-to-AWS migration with Terraform.

This was a solo build from first commit to production deployment: product design, UI/UX, frontend, backend, infrastructure, CI/CD, and operational runbooks.