← Back to portfolio

Data Engineering Case Study · Private Production Application

Job Board Monitor

A multi-source data application that turns heterogeneous job postings into an explainable, stateful decision workflow. It combines ingestion, normalization, classification, fit evaluation, deterministic prioritization, and application tracking in one Python and Streamlit system.

PythonStreamlitPlaywrightHTTP / APIspandasAutomated TestingGit

Finding a posting is easy. Maintaining decision context is not.

Job opportunities arrive through different career sites, applicant tracking systems, feeds, and saved links. Their fields, locations, titles, and descriptions are inconsistent. Repeated searches also create a state problem: which records are genuinely new, which have already been reviewed, and what decision was made last time?

I built Job Board Monitor to treat that workflow as a data-engineering problem. The system collects opportunities, converts them into a consistent model, preserves uncertainty, and maintains state between scans so each run adds information instead of recreating yesterday's inbox.

From heterogeneous sources to a controlled workflow

Each stage has one responsibility. Acquisition does not decide fit, classification does not hide eligibility uncertainty, and presentation does not become the source of record.

  1. 01
    Sources Career sites, ATS platforms, and job feeds
  2. 02
    Acquisition APIs, HTTP retrieval, and Playwright
  3. 03
    Normalization Shared fields, schema handling, and filtering
  4. 04
    Classification Role family, seniority, and geography
  5. 05
    Fit Evaluation Configurable weighted signals
  6. 06
    Priority Lane → fit → recency
  7. 07
    Persistent State Seen, saved, applied, and dismissed
  8. 08
    Streamlit UI Review and decision workflow
  9. 09
    Application Workflow A controlled handoff to the next action

Normalize first. Make decisions second.

01

Multi-source ingestion

Source adapters use APIs and direct HTTP retrieval where possible, with Playwright reserved for workflows that require browser behavior. The acquisition layer isolates source-specific logic before records enter the shared pipeline.

02

Normalization and filtering

Heterogeneous titles, locations, descriptions, identifiers, and URLs are transformed into a consistent representation. Validation and filtering happen before classification and scoring, giving downstream logic a stable contract.

03

Persistent monitoring

First-seen information and decision state distinguish new records from previously seen, saved, applied, or dismissed opportunities. Persistence is what turns periodic scraping into a useful monitoring system.

Two questions, answered by two separate layers

Question 1

What kind of opportunity is this?

Role classification establishes strategy before fit is considered.

  • PrimaryData Engineering
  • SecondaryAnalytics Engineering
  • SelectiveData Science
  • Off-strategyOther relevant roles

Question 2

How well does it align?

A configurable weighted-signal model evaluates fit within that strategic context. Ranking then follows an intentionally transparent rule:

Role lane→Fit score→Recency

This is deterministic decision logic, not machine-learning ranking. The design favors explainability over an opaque all-purpose score.

Preserve uncertainty and historical truth

Uncertainty remains visible

Potential state restrictions, California exclusions, relocation requirements, and conflicting remote or hybrid language become verification flags when eligibility is uncertain. Ambiguous classifications can remain unclassified rather than receiving a confident-looking guess.

History is not reconstructed

Saved and application records can retain the role family, priority lane, seniority, fit score, matched signals, and fit-signals version known at decision time. Unknown historical scores stay unknown rather than being recomputed under today's rules.

Tests protect decisions, not just functions

29 passing classification tests

Automated tests cover role-family classification, priority lanes, seniority, ambiguous titles, geographic restrictions, California eligibility, ML-heavy roles, historical backfill, and deterministic prioritization. The application-log and resume-artifact workflow has separate regression coverage.

This matters because a small rule change can silently reorder hundreds of opportunities or alter stored decision context. The test suite treats those behaviors as contracts.

A dated snapshot, not a product benchmark

Representative Scan September 2026
~400opportunities processed
132primary opportunities identified
132 Primary 66 Secondary 104 Selective 98 Off-strategy

Counts describe one representative scan after the latest classification work. No employer or application data is shown.

Document technical debt before optimizing it

Observed · Not yet normalized

The current weighted fit model can favor longer job descriptions because additional text creates more opportunities for signals to match. That bias is documented rather than presented as solved. I am collecting more evidence before adding normalization solely for sophistication, because a more complex score is not automatically a better decision tool.

One personal tool, several engineering disciplines

Data Engineering

Heterogeneous ingestion, normalization, schema handling, persistence, validation, and monitoring.

Analytics Engineering

Raw posting data transformed into stable, decision-oriented models with explicit meaning.

Software Engineering

Modular classification, configuration, state management, automated tests, and regression protection.

Applied Evaluation

Heuristic classification, weighted signals, ranking-behavior review, and iterative refinement without overstating ML.

Product Thinking

A real workflow problem improved through repeated use, observed failure modes, and explainable decisions.

Data Modeling

Business entities separated from their filesystem artifacts, with historical context preserved at decision time.

The architecture is public. The personal workflow is not.

This is a private production application. The operational repository and application data remain private because they contain personal workflow configuration. This case study focuses on the architecture and engineering decisions behind the system; it does not expose application history, private configuration, credentials, employer-specific search data, or a live production interface.