Back to Intika

Case study

KarmaDJ

Search and catalog engineering for a live karaoke platform.

KarmaDJ is a live karaoke request and catalog platform built around fast song discovery, catalog quality, and operational simplicity. Intika’s founder built it and runs it in real venues. The system evolved from watching karaoke nights: how people search on a phone in a loud room, what they fail to find, and what a venue needs from the admin side.

Product
KarmaDJ
Origin
Founder-built, in production
Domain
Live karaoke requests and catalog
Stack
PostgreSQLAlgoliaYouTube Data APIPlaywright

The system

Search and typeahead

Search runs on Algolia, with typeahead tuned for partial titles, artist names, and misspellings typed on a phone. PostgreSQL is the system of record. The index is derived from it and can be rebuilt from it at any time.

Content source and ingestion

Songs come from YouTube, restricted to a whitelist of karaoke channels so catalog quality is controlled at the source. Ingestion and catalog backfills run inside YouTube API quota limits, which shaped how often and how much of the catalog can be refreshed.

Title normalization

Karaoke video titles are inconsistent. Normalization strips channel boilerplate, tags, and formatting into a canonical title and artist. Alternate titles are kept, so a song is found under the names people actually use for it.

Coverage driven by real behavior

Searches that return nothing are logged. Those misses decide what gets added next, so the catalog grows where people are actually looking rather than by guesswork.

Admin tooling

Operators manage the catalog directly: title and artist corrections, alternate titles, channel whitelisting, and backfill runs. The tooling exists so catalog problems get fixed without an engineer.

Testing and production tradeoffs

Playwright covers the core flows in regression. Manual testing on real devices covers what automation misses in a venue. Managed search, one database, and scheduled backfills were chosen for cost and simplicity over doing everything in real time.

Architecture

SOURCE youtube whitelisted channels ingest backfill · quota-aware normalize titles · alternates catalog (pg) admin tools corrections · whitelist · runs algolia typeahead index app singers venues missed searches → next backfill playwright regression on core flows manual testing on real devices, in venues chosen for cost and simplicity: managed search · one database · scheduled backfills
Source, ingestion, catalog, and search path. Misses feed back into what gets added next.
missed searchlogged with the query reviewgrouped by how often it recurs sourcea whitelisted youtube channel backfillingest · normalize · index alternatesadded when people use other names checkthe same query, on a phone, in a venue
The catalog coverage loop. Real searches decide what gets added next.

What it taught us

Lesson

Search quality is a catalog problem before it is a search engine problem.

Lesson

Tooling for the operator is part of the system, not an extra.

Lesson

Simple and cheap holds up until real usage says otherwise.

Building or modernizing something with problems like these?

Start a project