← Back to portfolio

Data Pipeline Case Study · Applied Automation

Resume Screening Automation

A time-constrained Python workflow that transformed hundreds of unstructured PDF resumes into structured, reviewable candidate data. The system used explicit signals and deterministic scoring to support a hiring team—not machine learning and not an automated hiring decision.

PythonpdfplumberpandasregexCSVRule-based scoring

388 applications. Four days. No screening platform.

A hiring team needed to move from a folder of applicant PDFs to a defensible interview shortlist. Reading every document sequentially would consume the available window and make consistent comparisons difficult.

I built a focused data pipeline: extract text, structure relevant evidence, apply the same explicit rules to every record, and export results that a person could inspect before making a decision.

From documents to reviewable evidence

The design separated document processing from candidate evaluation. That made extraction failures, missing evidence, and scoring behavior visible instead of burying them in one opaque step.

  1. 01
    PDF Resumes388 unstructured applicant documents
  2. 02
    Text Extractionpdfplumber converts files into searchable text
  3. 03
    NormalizationConsistent casing, whitespace, and searchable fields
  4. 04
    Signal DetectionExplicit job-specific patterns and labels
  5. 05
    Weighted ScoringTransparent rules calibrated to the role
  6. 06
    Structured OutputReviewable CSV and candidate-level evidence
  7. 07
    Human ReviewShortlisting remains a hiring-team decision

Keep the system simple enough to inspect

01

Extract once, evaluate consistently

pdfplumber converted each PDF into text, which was then normalized into a common tabular representation. This separated format handling from the logic used to recognize qualifications.

02

Use role-specific signals

More than 20 signals were defined for the actual position rather than treating every keyword as equally useful. Programmatic labels made those matches available as structured columns.

03

Make scoring explainable

A weighted composite score summarized the configured evidence, but the underlying matched signals remained visible. Reviewers could understand why a record ranked where it did.

04

Preserve human judgment

The output supported review; it did not make the hiring decision. The shortlist could be checked against the extracted evidence, and unusual or incomplete documents could be handled manually.

The problem called for deterministic rules.

There was no labeled historical dataset that would justify training or validating a predictive model. The requirements were job-specific, the delivery window was short, and the hiring team needed to inspect the evidence behind each result.

A manual document queue became a structured review workflow.

388PDF resumes processed
20+job-specific signals evaluated
<4 daysavailable delivery window

These figures describe this specific stakeholder engagement, not a benchmark for other hiring workflows. The system produced structured CSV output and a reviewable shortlist for the hiring team.

Document extraction is not ground truth

Layout variance

Columns, tables, headers, and decorative elements can alter extraction order or separate related text.

Missing evidence

A signal not detected in extracted text does not prove that the candidate lacks the underlying experience.

Weighting choices

Scores reflect configured priorities. They are decision support, not an objective measurement of a person.

The responsible response is reviewability: retain extracted evidence, surface questionable documents, and keep a person accountable for the final interpretation.

Applied data engineering under a real deadline

Unstructured ingestionBatch PDF processing and text extraction
TransformationNormalized text converted into structured candidate records
Feature logicJob-specific evidence encoded as explicit signals
ExplainabilityWeights and matched evidence available for review
Delivery judgmentA deliberately simple solution matched to the deadline
Responsible automationHuman review retained for consequential decisions

The workflow is explained without publishing applicant data.

This case study does not include resumes, candidate names, contact details, raw extracted text, employer-specific hiring criteria, or individual scores. It documents the pipeline and its engineering decisions while keeping the underlying applicant information private.