•3 min read

I Got Tired of Tracking Job Applications, So I Built an ETL Pipeline

How I turned the frustrating process of job hunting into a live data engineering pipeline using Python, Gemini, and Streamlit.

Data EngineeringPythonETLStreamlit

Job searching is a strange full-time job that does not come with a salary. One minute you are confidently applying for a role, and the next, you are lost in a sea of portal logins, recruiter emails, and duplicate CV uploads.

My tracking spreadsheet quickly became part of the problem: a graveyard of missing dates, outdated statuses, and forgotten follow-ups. As someone who works with data, I had to admit the truth. My job search had become a data engineering problem. Instead of colour-coding my spreadsheet again, I built an automated ETL pipeline.

The Architecture of Job Search HQ The information I needed was already in my inbox, buried beneath banners and disclaimers. I needed a system to read these unstructured messages and convert them into consistent records. The resulting pipeline runs three times daily via GitHub Actions:

Extract: The Gmail API retrieves relevant application emails, interview invites, and recruiter updates.

Transform: Python preprocesses the text and sends it to Google Gemini to extract a structured schema (company, salary, role, and deadlines).

Load: The system validates the data, checks for duplicate company and title identifiers, and loads the clean records into Google Sheets.

Turning Chaos Into Intelligence Once the records are structured, a live Streamlit dashboard turns them into operational intelligence. It tracks my recruitment funnel from resume screening to offers, helping me identify exactly where the process slows down.

Source Performance Analysis: It separates platforms that generate high application volume from those that actually lead to interviews.

The "Needs Attention" Queue: It highlights overdue deadlines, upcoming interviews, and active applications sitting dormant for over seven days.

Role Categorization: It groups job titles (such as Data Analytics, AI/LLM, and Business Intelligence) to compare my target roles against my actual application habits.

Engineering the Messy Details The final dashboard is clean, but the incoming data is chaotic. I had to engineer a mixed-date parser to handle conflicting ISO formats, day-first strings, and spreadsheet serial numbers. Missing information, like unlisted salaries or ambiguous locations, is caught by the cleaning layer and properly normalized as null values rather than breaking the analytics models.

What I Learned This project started because I hated updating a spreadsheet, but it became a masterclass in API integration, LLM-assisted extraction, and operational analytics. Without data, every unanswered application feels like a mystery. With structured telemetry, the job hunt becomes a measurable, optimizable pipeline.

I am still a candidate looking for the right opportunity, but now I have a personalized analytics department running the search.

Explore the Project: