Generated: 2026-06-19 Audited by: AntiGravity architectural analysis
JobHunter (codenamed Arachnode) is an event-driven microservice system designed to automate the top-of-funnel pipeline for software engineering job applications. It crawls startup directories and job platforms, deduplicates job postings using a Redis stream and a PostgreSQL database, performs OSINT-based contact discovery using the GitHub API and LinkedIn search, and automatically generates personalized cold outreach emails via Jinja2 and local LLMs (Ollama). The system solves the problem of manual tracking, candidate discovery, and boilerplate email drafting for bulk job searching.
Currently, the codebase is in a highly functional but partially hardened state. The core orchestration, including the Redis stream pub/sub mechanism, PostgreSQL data persistence, API gateway, and individual services for scraping and email generation, are fully built. There is a background scheduler that automates sweeps. New features like resume parsing, an Unstop scraper, weekly email digests, and follow-up email logic have been added beyond the initial plan. However, some aspects remain partially built or incomplete: there is a critical bug where the scheduler tries to run Scrapy spiders in a subprocess without having scrapy installed, contact discovery Playwright binaries are improperly cached in Docker, and there is zero authentication protecting the exposed API endpoints.
Overall, the architecture demonstrates a mature, scalable decoupling of concerns through bounded contexts and asynchronous event handling. The code quality is generally high, utilizing Python 3.11 features, Pydantic validation, and async I/O. However, the system is not entirely ready for open-source contributors without addressing the missing end-to-end tests, Docker dependency bugs, and the complete lack of API authentication, which poses a severe security risk if deployed to a public VPS.
jobCrawler/
βββ .claude/ # Claude agent instruction files β COMPLETE
β βββ agents/ # 24 agent profiles for different QA/auditing tasks β COMPLETE
βββ agent-search-service/ # Intent-driven search engine & ATS integration β COMPLETE
β βββ ats_client.py # Lever, Greenhouse, Ashby ATS connectors β COMPLETE
β βββ Dockerfile # Multi-stage Docker build β COMPLETE
β βββ main.py # Search session routes & SSE streaming β COMPLETE
β βββ requirements.txt # Python dependencies β COMPLETE
β βββ search_executor.py # DuckDuckGo and ATS search runner β COMPLETE
βββ aggregator-service/ # Jobs aggregator and deduplicator β COMPLETE
β βββ consumer.py # Background Redis stream consumer β COMPLETE
β βββ db.py # asyncpg PostgreSQL adapter β COMPLETE
β βββ Dockerfile # Multi-stage Docker build β COMPLETE
β βββ main.py # FastAPI endpoints for jobs and stats β COMPLETE
β βββ matcher.py # Semantic ranking utility β COMPLETE
β βββ models.py # Pydantic schemas β COMPLETE
β βββ requirements.txt # Python dependencies β COMPLETE
β βββ test_matcher.py # Unit tests for matcher β COMPLETE
β βββ utils/ # Helper utilities (date_utils.py) β COMPLETE
βββ contact-discovery-service/ # OSINT contact discovery β COMPLETE
β βββ discovery.py # Pipeline: domains, emails, GitHub & LinkedIn scraping β COMPLETE
β βββ Dockerfile # Docker build (Buggy Playwright cache) β PARTIAL
β βββ main.py # FastAPI endpoints β COMPLETE
β βββ requirements.txt # Python dependencies β COMPLETE
β βββ storage.py # asyncpg PostgreSQL adapter β COMPLETE
β βββ verifier.py # SMTP validation with rate limits β COMPLETE
βββ crawler-service/ # Scrapy-based web crawlers β COMPLETE
β βββ crawler/
β β βββ spiders/ # Scrapy spiders
β β β βββ base_spider.py # Base startup spider class β COMPLETE
β β β βββ cutshort_spider.py # Cutshort spider β COMPLETE
β β β βββ github_org_spider.py # GitHub Org spider β COMPLETE
β β β βββ glassdoor.py # Glassdoor spider β COMPLETE
β β β βββ remotive_spider.py # Remotive spider β COMPLETE
β β β βββ wellfound_spider.py # Wellfound spider β COMPLETE
β β β βββ yc_spider.py # YC Jobs spider β COMPLETE
β βββ Dockerfile # Docker build (Runs as root) β PARTIAL
β βββ read_stream.py # Helper to read from Redis stream β COMPLETE
β βββ run_local.sh # Bash script for local execution β COMPLETE
β βββ scrapy.cfg # Scrapy configuration β COMPLETE
β βββ tests/ # Tests for the crawler β PARTIAL (Only ATS detector)
βββ email-generator-service/ # Cold email draft generator β COMPLETE
β βββ Dockerfile # Multi-stage Docker build β COMPLETE
β βββ fallbacks.yaml # Fallback text strings β COMPLETE
β βββ generator.py # Email generation orchestration β COMPLETE
β βββ generator_digest.py # Weekly digest generation β COMPLETE
β βββ mailer.py # Gmail SMTP sender β COMPLETE
β βββ main.py # FastAPI endpoints (resume upload, digest) β COMPLETE
β βββ ollama_client.py # Local LLM wrapper β COMPLETE
β βββ resume_parser.py # Resume parsing logic β COMPLETE
β βββ storage.py # asyncpg PostgreSQL adapter β COMPLETE
β βββ templates/ # Jinja2 templates for emails β COMPLETE
β βββ test_resume_parser.py # Unit tests for resume parser β COMPLETE
β βββ RESUME_PARSER_EXAMPLES.md # Examples of personalized emails β COMPLETE
βββ gateway/ # Unified API Gateway and Proxy β COMPLETE
β βββ dashboard.html # Single-page Vanilla JS Dashboard β COMPLETE
β βββ Dockerfile # Docker build β COMPLETE
β βββ main.py # FastAPI router fanout and workflows β COMPLETE
β βββ proxy.py # httpx request forwarding β COMPLETE
βββ gemini-agent-service/ # Gemini AI assistant service β COMPLETE
β βββ agent.py # Gemini LLM orchestrator β COMPLETE
β βββ Dockerfile # Multi-stage Docker build β COMPLETE
β βββ main.py # FastAPI endpoints β COMPLETE
β βββ requirements.txt # Python dependencies β COMPLETE
β βββ storage.py # In-memory session history storage β COMPLETE
β βββ tools.py # Search, resume, and email drafting tools β COMPLETE
βββ scheduler/ # APScheduler automation pipelines β COMPLETE
β βββ Dockerfile # Docker build (Missing scrapy dependency) β PARTIAL
β βββ logger.py # Custom JSON logger β COMPLETE
β βββ main.py # APScheduler background daemon β COMPLETE
β βββ tasks.py # Tasks for scrape, discover, draft, digest, followup β COMPLETE
βββ scraper-service/ # On-demand Playwright scrapers β COMPLETE
β βββ discovery/ # Dork builder utilities β COMPLETE
β βββ Dockerfile # Multi-stage Docker build (Correct Playwright cache) β COMPLETE
β βββ emit.py # Shared Redis stream emitter β COMPLETE
β βββ main.py # FastAPI endpoints β COMPLETE
β βββ scrapers/ # Platform scrapers
β β βββ base.py # Base scraper class β COMPLETE
β β βββ google_dork.py # Google Dork scraper β COMPLETE
β β βββ internshala.py # Internshala scraper β COMPLETE
β β βββ linkedin.py # LinkedIn scraper β COMPLETE
β β βββ naukri.py # Naukri scraper β COMPLETE
β β βββ unstop.py # Unstop scraper β COMPLETE
β βββ tests/ # Unit tests β PARTIAL
β βββ UNSTOP.md # Documentation for Unstop scraper β COMPLETE
βββ tests/ # Cross-service tests β PARTIAL
β βββ contract/ # JSON schema validation tests β COMPLETE
β βββ integration/ # Redis and Postgres integration tests β COMPLETE
β βββ unit/ # Independent service logic tests β COMPLETE
β βββ conftest.py # Pytest fixtures β COMPLETE
βββ workflows/ # GitHub Issue/PR templates β COMPLETE
βββ docker-compose.yml # Infrastructure orchestrator β COMPLETE
βββ README.md # Comprehensive documentation β COMPLETE
- Status: COMPLETE
- Language and framework: Python / Scrapy
- Port: None (Runs as a one-shot process)
- Entry point file:
crawler-service/crawler/spiders/*(via Scrapy CLI) - What it does: Navigates startup directories (YC, Remotive, Wellfound, etc.), extracts job listings using XPath/CSS selectors, and publishes items as JSON payloads to the
jobs:rawRedis stream. - What is working: Successfully scrapes flat HTML and emits structured data.
- What is broken or incomplete: Fails to run automatically inside the Scheduler container due to a missing Scrapy dependency and inaccessible project paths. Container runs as root.
- External dependencies it calls: Redis (publishes to stream).
- What calls it: Called via the
docker runequivalent or as a subprocess by the Scheduler (currently broken). - Known issues or code smells spotted during audit: It runs as
rootin Docker. Test coverage is nearly non-existent.
- Status: COMPLETE
- Language and framework: Python / FastAPI / Playwright
- Port: 8001
- Entry point file:
scraper-service/main.py - What it does: Handles on-demand JavaScript-heavy browser scraping for platforms like LinkedIn, Naukri, Internshala, and Unstop. Runs them concurrently via a BackgroundTask and emits normalized job entities to the Redis stream.
- What is working: The FastAPI endpoints and Playwright scrapers function as intended, successfully emitting jobs to Redis.
- What is broken or incomplete: Google Dorks discovery is implemented but noted as "demo-friendly" without emitting jobs.
- External dependencies it calls: Redis (publishes to stream).
- What calls it: API Gateway (
POST /api/scrape). - Known issues or code smells spotted during audit: Playwright runs with
--no-sandbox.
- Status: COMPLETE
- Language and framework: Python / FastAPI / asyncpg / redis-py
- Port: 8000
- Entry point file:
aggregator-service/main.py - What it does: Runs a background asyncio consumer loop to read from the
jobs:rawRedis stream, deduplicates jobs using MD5 hashes of the normalized company and role, and persists them into PostgreSQL. Exposes a queryable REST API for job analytics and listings. - What is working: Redis consumer group mechanics, database idempotent insertions, and query filtering.
- What is broken or incomplete: Semantic ranking (resume parsing) logic is handled on read (
GET /jobs?resume=), which could become slow on large data sets since it recalculates ranks on the fly. - External dependencies it calls: Redis (Stream reading), PostgreSQL (CRUD operations).
- What calls it: API Gateway (routes
/api/jobs/*). - Known issues or code smells spotted during audit: Missing pagination metadata in the API response (returns a flat list up to
limit).
- Status: COMPLETE
- Language and framework: Python / FastAPI / httpx / Playwright
- Port: 8002
- Entry point file:
contact-discovery-service/main.py - What it does: Uses OSINT techniques to find recruiter and engineering manager contacts for a specific company. It infers domains via Clearbit, detects email patterns via GitHub commit logs, scrapes names from LinkedIn and GitHub orgs, and validates emails via SMTP probes.
- What is working: The entire pipeline logic, rate-limited SMTP verification, and asynchronous PostgreSQL persistence.
- What is broken or incomplete: The Dockerfile runs
playwright installas root before theUSER appuserdirective, causing the Playwright binary to be placed in an inaccessible cache folder for the runtime user. - External dependencies it calls: PostgreSQL, Clearbit Autocomplete API, GitHub API, LinkedIn, arbitrary SMTP servers.
- What calls it: API Gateway (
POST /api/discoverand/api/workflow/apply). - Known issues or code smells spotted during audit: LinkedIn scraping is highly susceptible to authwalls. Rate limit dictionary (
_domain_rate) is in-memory, meaning it won't sync if scaled horizontally.
- Status: COMPLETE
- Language and framework: Python / FastAPI / Ollama / Jinja2
- Port: 8003
- Entry point file:
email-generator-service/main.py - What it does: Evaluates Jinja2 templates, interacts with a local Ollama instance for LLM-powered context mapping, parses PDF/TXT resumes to build candidate context, drafts personalized cold emails, stores drafts in PostgreSQL, and sends emails via Gmail SMTP.
- What is working: Resume parsing, email generation, template rendering, and SMTP dispatch.
- What is broken or incomplete: Currently only supports sending via a hardcoded
GMAIL_ADDRESSandGMAIL_APP_PASSWORD. - External dependencies it calls: PostgreSQL, Ollama (local/remote), Gmail SMTP (port 465).
- What calls it: API Gateway (
POST /api/generate,/api/emails/*,/api/workflow/apply,/api/digest). - Known issues or code smells spotted during audit: The
POST /resumeendpoint takes a file upload directly but has no authorization, allowing arbitrary file uploads (though limited to 5MB and not written to disk).
- Status: COMPLETE
- Language and framework: Python / FastAPI / httpx / Vanilla JS
- Port: 8080
- Entry point file:
gateway/main.py - What it does: Acts as the unified public proxy for all internal microservices. It forwards requests via
httpx, hosts the staticdashboard.htmlsingle-page application, and manages composite endpoints. - What is working: Request proxying, dashboard serving, API key authentication (
APIKeyMiddleware), and cross-service orchestration. - What is broken or incomplete: None. API key auth defaults to unauthenticated in local dev mode if
JOBHUNTER_API_KEYis omitted. - External dependencies it calls: Aggregator, Scraper, Contact Discovery, Email Generator, Agent Search, and Gemini Agent services.
- What calls it: User via Web Browser, Scheduler Service.
- Status: COMPLETE
- Language and framework: Python / FastAPI / httpx / DuckDuckGo
- Port: 8009
- Entry point file:
agent-search-service/main.py - What it does: On-demand intent-driven job search engine. Interprets natural language queries, performs Google dorking via DuckDuckGo, polls public ATS APIs (Lever, Greenhouse, Ashby), enforces freshness filtering, and streams results over SSE.
- What is working: SSE streaming, ATS client indexing, freshness filtering, and relevance scoring.
- External dependencies it calls: DuckDuckGo Web API, Lever, Greenhouse, Ashby ATS endpoints, Redis Stream (
jobs:raw). - What calls it: Gateway (
/api/agent/search).
- Status: COMPLETE
- Language and framework: Python / FastAPI / google-generativeai
- Port: 8010
- Entry point file:
gemini-agent-service/main.py - What it does: AI assistant powered by Gemini 2.0. Provides interactive chat, resume tuning against job descriptions, and custom cold email drafting.
- What is working: Conversational AI chat, resume enhancer, cold email drafter, session history, and graceful degradation if
GEMINI_API_KEYis missing. - External dependencies it calls: Google Gemini API (
gemini-2.0-flash). - What calls it: Gateway (
/api/agent/*).
- Status: PARTIAL
- Language and framework: Python / APScheduler / httpx
- Port: None (Background daemon)
- Entry point file:
scheduler/main.py - What it does: Runs timed automation sweeps. Every 8 hours it triggers scrapers and local Scrapy spiders. Every 24 hours it discovers contacts for new jobs and drafts emails. Weekly, it sends an email digest. Daily, it drafts follow-ups.
- What is working: Job scheduling, offset execution, weekly digests, and follow-ups.
- What is broken or incomplete: The
run_scrape_cycleattempts to executesubprocess.run(["scrapy", "crawl", spider])inside the scheduler container. However,scrapyis not installed in the scheduler'srequirements.txt, and the crawler directory is not mounted by default. - External dependencies it calls: Gateway Service (
GET/POSTvia HTTP). - What calls it: Self-triggered based on chron intervals.
- Known issues or code smells spotted during audit: Subprocess execution is an anti-pattern in Dockerized microservices. The scheduler should instead trigger the Crawler container or expose an endpoint on the crawler.
CREATE EXTENSION IF NOT EXISTS "pgcrypto";
CREATE TABLE IF NOT EXISTS jobs (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
company TEXT NOT NULL,
role TEXT NOT NULL,
source TEXT,
url TEXT,
stack TEXT[],
product TEXT,
location TEXT,
posted_at TIMESTAMPTZ,
status TEXT NOT NULL DEFAULT 'new',
created_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
);
CREATE TABLE IF NOT EXISTS contacts (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
job_id UUID REFERENCES jobs(id) ON DELETE SET NULL,
company TEXT NOT NULL,
domain TEXT,
name TEXT,
email TEXT,
role TEXT,
source TEXT,
verified TEXT NOT NULL DEFAULT 'unverified',
created_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
CONSTRAINT contacts_company_email_key UNIQUE (company, email)
);
CREATE TABLE IF NOT EXISTS emails (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
job_id UUID REFERENCES jobs(id) ON DELETE SET NULL,
contact_id UUID REFERENCES contacts(id) ON DELETE SET NULL,
template TEXT NOT NULL,
subject TEXT NOT NULL,
body TEXT NOT NULL,
generated_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
sent_at TIMESTAMPTZ,
status TEXT NOT NULL DEFAULT 'draft'
);
-- Indexes
CREATE INDEX IF NOT EXISTS idx_jobs_stack ON jobs USING GIN (stack);
CREATE INDEX IF NOT EXISTS idx_jobs_posted_at ON jobs (posted_at DESC NULLS LAST);
CREATE INDEX IF NOT EXISTS idx_jobs_status ON jobs (status);
CREATE INDEX IF NOT EXISTS idx_contacts_company ON contacts (company);
CREATE INDEX IF NOT EXISTS idx_contacts_job_id ON contacts (job_id) WHERE job_id IS NOT NULL;
CREATE INDEX IF NOT EXISTS idx_contacts_email ON contacts (email) WHERE email IS NOT NULL;
CREATE INDEX IF NOT EXISTS idx_emails_job_id ON emails (job_id) WHERE job_id IS NOT NULL;
CREATE INDEX IF NOT EXISTS idx_emails_contact_id ON emails (contact_id) WHERE contact_id IS NOT NULL;
CREATE INDEX IF NOT EXISTS idx_emails_status ON emails (status);ββββββββββββββββββββββββββ
β jobs β
ββββββββββββββββββββββββββ€
β id (UUID) [PK] βββββββββββ
β company (TEXT) β β
β role (TEXT) β β
β source (TEXT) β β
β url (TEXT) β β
β stack (TEXT[]) β β
β product (TEXT) β β
β location (TEXT) β β
β posted_at (TIMESTAMPTZ)β β
β status (TEXT) βββββββ β
β created_at (TIMESTAMPTZβ β β
ββββββββββββ¬ββββββββββββββ β β
β β β
β β β
ββββββββββββ΄ββββββββββββββ β β
β contacts β β β
ββββββββββββββββββββββββββ€ β β
β id (UUID) [PK] ββββ β β
β job_id (UUID) [FK] ββββ β β
β company (TEXT) β β β
β domain (TEXT) β β β
β name (TEXT) β β β
β email (TEXT) β β β
β role (TEXT) β β β
β source (TEXT) β β β
β verified (TEXT) β β β
β created_at (TIMESTAMPTZβ β β
ββββββββββββ¬ββββββββββββββ β β
β β β
β β β
ββββββββββββ΄ββββββββββββββ β β
β emails β β β
ββββββββββββββββββββββββββ€ β β
β id (UUID) [PK] β β β
β job_id (UUID) [FK] βββββββββββ
β contact_id (UUID) [FK] βββββββ
β template (TEXT) β
β subject (TEXT) β
β body (TEXT) β
β generated_at(TIMESTAMP)β
β sent_at (TIMESTAMPTZ) β
β status (TEXT) β
ββββββββββββββββββββββββββ
- Are indexes appropriate for the query patterns in the code? Yes, GIN indexes are efficiently utilized for array overlap operations on the
stackarray. Partial indexes onjob_id,contact_id, andemailappropriately map the API access patterns. - Are there missing indexes that would cause slow queries? There is no index on
jobs.companydespite contact discovery frequently filtering by it. Contact discovery queries useWHERE company ILIKE $1, which scans the entire table regardless of an index, but an exact match query oncompanywould benefit from an index. - Are there any N+1 query patterns in the codebase? Yes. The composite
POST /api/workflow/applyendpoint runs distinct queries per service rather than taking advantage of relational joins, acting as a serialized N+1 over HTTP. - Are foreign key constraints enforced or just implied? They are explicitly enforced at the database level using
ON DELETE SET NULL. - What schema changes have been made since the original design? No migration tool (like Alembic) is used. The schema is defined as idempotent DDL in
db.pyandstorage.pyacross services.
- Service: Gateway (fans out to all services)
- Status: WORKING
- Purpose: Provides a complete system liveness check.
- Request body: None
- Query parameters: None
- Response schema:
{"gateway": "ok", "services": [...]} - Calls: Aggregator
/health, Scraper/health, Contact/health, Email-Gen/health. - Known issues: Returns 207 Multi-Status if any internal service is down, which is good API design, but it will block synchronously waiting for unresponsive services to timeout.
- Service: Gateway
- Status: WORKING
- Purpose: Returns the most recent JSON run summary produced by the scheduler.
- Request body: None
- Query parameters: None
- Response schema: JSON file payload from disk.
- Calls: Local filesystem (
/data/run_summary.json). - Known issues: Will crash if the file is locked or malformed since there is no file locking mechanism between the Scheduler and the Gateway.
- Service: Aggregator
- Status: WORKING
- Purpose: List and filter jobs.
- Request body: None
- Query parameters:
role(str),stack(str),status(str),sort(str: 'latest', 'oldest'),limit(int),resume(str). - Response schema:
List[JobOut] - Calls: PostgreSQL
jobstable. - Known issues: The
resumeparameter triggers a blockingrank_jobs(jobs, resume)semantic sorting pass, which can stall the async loop.
- Service: Aggregator
- Status: WORKING
- Purpose: Export filtered jobs as CSV.
- Request body: None
- Query parameters:
role(str),stack(str),status(str),sort(str),format(str). - Response schema:
StreamingResponse(CSV text) - Calls: PostgreSQL
jobstable. - Known issues: None. Efficient chunked iteration is used.
- Service: Aggregator
- Status: WORKING
- Purpose: Fetch a single job.
- Request body: None
- Query parameters: None
- Response schema:
JobOut - Calls: PostgreSQL
jobstable. - Known issues: None.
- Service: Aggregator
- Status: WORKING
- Purpose: Update job application status.
- Request body:
{"status": "new|applied|ignored"} - Query parameters: None
- Response schema:
JobOut - Calls: PostgreSQL
jobstable. - Known issues: None.
- Service: Aggregator
- Status: WORKING
- Purpose: Aggregate counts by source and status.
- Request body: None
- Query parameters: None
- Response schema:
StatsOut - Calls: PostgreSQL
jobstable. - Known issues: Uses
COUNT(*)over the entire table without time bounds, which will become slow at scale.
- Service: Scraper
- Status: WORKING
- Purpose: Trigger all scraping scripts concurrently.
- Request body:
{"role": "string", "stack": ["string"]} - Query parameters: None
- Response schema:
ScrapeResponse - Calls: Playwright browser scripts, Redis Stream (
jobs:raw). - Known issues: Runs as a FastAPI BackgroundTask but returns immediately; the caller receives no feedback on actual success or failure.
- Service: Contact
- Status: WORKING
- Purpose: Find contacts for a company.
- Request body:
{"company": "string", "job_id": "UUID", "roles": ["string"], "domain": "string"} - Query parameters: None
- Response schema:
{"triggered": true, "company": "...", "message": "..."} - Calls: Clearbit API, GitHub API, LinkedIn, arbitrary SMTP servers, PostgreSQL
contactstable. - Known issues: Extensive third-party calls executed in the background without retry logic.
- Service: Contact
- Status: WORKING
- Purpose: List contacts for a company.
- Request body: None
- Query parameters:
company(str, required) - Response schema:
List[ContactOut] - Calls: PostgreSQL
contactstable. - Known issues: Filters via
ILIKE %company%which causes a full table scan.
- Service: Contact
- Status: WORKING
- Purpose: List contacts associated directly with a job.
- Request body: None
- Query parameters: None
- Response schema:
List[ContactOut] - Calls: PostgreSQL
contactstable. - Known issues: None.
- Service: Contact
- Status: WORKING
- Purpose: Delete a specific contact.
- Request body: None
- Query parameters: None
- Response schema: 204 No Content
- Calls: PostgreSQL
contactstable. - Known issues: None.
- Service: Email-Gen
- Status: WORKING
- Purpose: Generate a personalized email.
- Request body:
GenerateRequest(template, candidate parameters, job/contact ids). - Query parameters: None
- Response schema:
GenerateResponse - Calls: PostgreSQL
jobsandcontactstables, Ollama REST API. - Known issues: Evaluates LLM completion directly inside the request/response lifecycle. Long inference times will cause API timeouts.
- Service: Email-Gen
- Status: WORKING
- Purpose: List emails for a given job.
- Request body: None
- Query parameters:
job_id(UUID, required) - Response schema:
List[EmailOut] - Calls: PostgreSQL
emailstable. - Known issues: None.
- Service: Email-Gen
- Status: WORKING
- Purpose: Fetch a specific email.
- Request body: None
- Query parameters: None
- Response schema:
EmailOut - Calls: PostgreSQL
emailstable. - Known issues: None.
- Service: Email-Gen
- Status: WORKING
- Purpose: Manually update an email's status.
- Request body:
{"status": "draft|sent|replied"} - Query parameters: None
- Response schema:
EmailOut - Calls: PostgreSQL
emailstable. - Known issues: None.
- Service: Email-Gen
- Status: WORKING
- Purpose: Send a generated email via Gmail.
- Request body: None
- Query parameters: None
- Response schema:
EmailOut - Calls: PostgreSQL
contactstable, Gmail SMTP server. - Known issues: Blocking operation executed in thread-pool.
- Service: Gateway
- Status: WORKING
- Purpose: Execute the discovery-to-draft orchestration synchronously.
- Request body:
{"job_id": "UUID", "template": "...", "roles": ["..."]} - Query parameters: None
- Response schema:
{"job": {...}, "contacts": [...], "draft_email": {...}} - Calls: Aggregator GET, Contact POST, Contact GET, Email-Gen POST.
- Known issues: Implements an arbitrary
asyncio.sleep(3)to wait for background discovery. Highly fragile and prone to race conditions if discovery takes longer than 3 seconds.
| Method | Path | Service | Status |
|---|---|---|---|
| GET | /api/health | Gateway | WORKING |
| GET | /api/summary | Gateway | WORKING |
| GET | /api/jobs | Aggregator | WORKING |
| GET | /api/jobs/export | Aggregator | WORKING |
| GET | /api/jobs/{id} | Aggregator | WORKING |
| PATCH | /api/jobs/{id}/status | Aggregator | WORKING |
| GET | /api/stats | Aggregator | WORKING |
| POST | /api/scrape | Scraper | WORKING |
| POST | /api/discover | Contact | WORKING |
| GET | /api/contacts | Contact | WORKING |
| GET | /api/contacts/{id} | Contact | WORKING |
| DELETE | /api/contacts/{id} | Contact | WORKING |
| POST | /api/generate | Email-Gen | WORKING |
| GET | /api/emails | Email-Gen | WORKING |
| GET | /api/emails/{id} | Email-Gen | WORKING |
| PATCH | /api/emails/{id}/status | Email-Gen | WORKING |
| POST | /api/emails/{id}/send | Email-Gen | WORKING |
| POST | /api/agent/search | Agent-Search | WORKING |
| POST | /api/agent/chat | Gemini-Agent | WORKING |
| POST | /api/agent/enhance-resume | Gemini-Agent | WORKING |
| POST | /api/agent/draft-email | Gemini-Agent | WORKING |
| POST | /api/workflow/apply | Gateway | WORKING |
- Scheduler triggers the crawler at specified intervals (or user triggers via
POST /api/scrape). - Playwright and Scrapy scrapers fetch platforms, parse HTML, and build an intermediary Python dictionary.
- The scraper's
emit.pynormalizes fields and publishes JSON payloads to thejobs:rawRedis stream viaXADD. - The Aggregator service runs an
XREADGROUPconsumer loop. It reads the stream, normalizes the company and role to generate an MD5 deduplication hash (dedup:agg:hash). - If the hash doesn't exist in Redis, the Aggregator saves the job to the PostgreSQL
jobstable andXACKs the message. - The Scheduler executes a discovery sweep, calling
POST /api/discoverfor new jobs. - The Contact service receives the job, finds contacts, validates their emails, and stores them in the
contactstable (linked viajob_id). - The Scheduler executes a drafting sweep, calling
POST /api/generate. - The Email service fetches candidate context, requests Ollama LLM completion, maps Jinja2 templates, and inserts a draft into the
emailstable. - The user clicks "Send" on the Dashboard, invoking
POST /api/emails/{id}/send, which fires an SMTP request to Gmail and updates the DB status to 'sent'.
- Stream name(s) found in the code:
jobs:raw - Producer services and what they emit:
scraper-serviceandcrawler-serviceemit normalized dictionary representations of job postings. - Consumer services and what they do with messages:
aggregator-serviceverifies duplication logic via MD5 hash lookups on Redis and inserts non-duplicates into PostgreSQL. - Consumer group configuration: Group
aggregator-group, consumeraggregator-1. UsesXAUTOCLAIMto recover failed messages. - Current maxlen setting and whether it is appropriate:
maxlen=50_000approximate. Appropriate for a personal data funnel. - Any message loss risk identified: Low. Messages are only
XACKed after successful Postgres insertion.
| Field | Crawler output | Redis Stream | After aggregator | API response |
|---|---|---|---|---|
| id | N/A | N/A | UUID | UUID |
| company | str | JSON string | TEXT | str |
| role | str | JSON string | TEXT | str |
| source | str | JSON string | TEXT | str |
| url | str | JSON string | TEXT | str |
| stack | list[str] | JSON string array | TEXT[] | list[str] |
| product | str | JSON string | TEXT | str |
| location | str | JSON string | TEXT | str |
| posted_at | None / str | JSON string | TIMESTAMPTZ | str (ISO format) |
| status | N/A | N/A | TEXT | str ('new'/'applied'/'ignored') |
| created_at | N/A | N/A | TIMESTAMPTZ | str (ISO format) |
- Added by: Unknown contributor.
- What it does: Scrapes both
/jobsand/internshipson unstop.com. Uses Playwright to render the Angular SPA. - Files changed or added:
scraper-service/scrapers/unstop.py,scraper-service/run_unstop.py,scraper-service/UNSTOP.md,tests/unit/test_unstop_parser.py,scraper-service/main.py - How it integrates with the existing architecture: Integrated into the
/scrapeendpoint concurrently with other Playwright scrapers. - Test coverage: Yes (Unit tests in
tests/unit/test_unstop_parser.py) - Documentation: Yes (
UNSTOP.mdprovided).
- Added by: Unknown contributor (Planned feature completed).
- What it does: Accepts PDF/TXT file uploads, parses candidate context (skills, experience, role), and pipes context into Ollama/Jinja2 to hyper-personalize generated drafts.
- Files changed or added:
email-generator-service/resume_parser.py,email-generator-service/main.py,email-generator-service/test_resume_parser.py,email-generator-service/RESUME_PARSER_EXAMPLES.md - How it integrates with the existing architecture: Exposed via
POST /resumeendpoint to fetch JSON context. Data is passed intoPOST /generate. - Test coverage: Yes (
test_resume_parser.py) - Documentation: Yes (
RESUME_PARSER_EXAMPLES.md).
- Added by: Unknown contributor.
- What it does: Computes a week label and drafts a weekly email summary of discovered jobs. Sent out every Sunday via APScheduler.
- Files changed or added:
email-generator-service/generator_digest.py,scheduler/tasks.py(addedrun_digest_cycle),scheduler/main.py - How it integrates with the existing architecture: Exposed via
POST /digeston the Email service. Triggers SMTP immediately without saving drafts to the database. - Test coverage: Partial
- Documentation: No
- Added by: Unknown contributor.
- What it does: Automatically checks the database for emails sent >
FOLLOWUP_DAYS(7 days) ago and drafts follow-up templates if no replies were noted. - Files changed or added:
scheduler/tasks.py(addedrun_followup_cycle),scheduler/main.py - How it integrates with the existing architecture: Uses existing
/api/emailsand/api/generateendpoints via Gateway API. - Test coverage: No
- Documentation: No
- Services added beyond the original plan: None explicitly, though the
.claude/agents/directory indicates Claude AI agents are heavily integrated as an orchestration layer for codebase monitoring, auditing, and maintenance. - Services that were merged or split: The original 7-service design is still fully intact.
- Architectural patterns that changed: None major. The codebase strictly adhered to the REST API / Redis event stream duality.
- New infrastructure components introduced: None.
- Any architectural debt introduced: The Scheduler service executes
subprocess.run(["scrapy", "crawl", spider]). This heavily couples the daemon scheduler to the crawler binaries. Furthermore, thedocker-compose.ymlmounts do not support this, so the container errors out entirely when it attempts to run Scrapy.
ββββββββββββββββββββββββββ
βββββββββββββββββββββββ Jobs via REST POST β β
β ββββββββββββββββββββββββββββΊβ Gateway Service β
β Scheduler Service β β (:8080) β
β (APScheduler) β Trigger operations β β
ββββββββββββ¬βββββββββββ ββββββββ¬βββββββββ¬βββββββββ
β β β
β Trigger spiders via POST REST β β REST
β β β
ββββββββββββΌβββββββββββ ββββββββΌβββββββββΌβββββββββ
β β Jobs via POST β β
β Platform Scraper ββββββββββββββββββββββββββββΊβ Aggregator Service β
β (:8001) β β (:8000) β
βββββββββββββββββββββββ ββββββββ¬ββββββββββββββββββ
β β²
βββββββββββββββββββββββ Jobs via Stream β β Store
β β ββββββββΌβββββββββΌβββββββββ
β Crawler Service βββββββ[ Redis ]ββββββββββββΊβ PostgreSQL β
β (Scrapy spider) β Stream β Database DB β
βββββββββββββββββββββββ ββββββββ¬βββββββββ¬βββββββββ
β β
βββββββββββββββββββββββ ββββββββΌβββββββββΌβββββββββ
β β Trigger via REST POST β β
β Email Service βββββββββββββββββββββββββββββ€ Contact Discovery β
β (:8003) β β (:8002) β
βββββββββ¬ββββββββββββββ ββββββββββββββββββββββββββ
β
βββββββββΌββββββββββββββ
β Ollama Local LLM / β
β Gmail SMTP Gateway β
βββββββββββββββββββββββ
- Aggregator Service: Standard Python 3.11 slim.
appuserimplemented correctly. - Contact Discovery Service: Implements Playwright.
appuseris implemented afterplaywright install chromiumwithout explicit path declarations. Binaries are locked out of the runtime user's accessibility. - Crawler Service: Does not implement
appuserat all. Runs as root. Unnecessarywget,curl,gnupgpackages installed for Chrome but Playwright handles this cleanly. - Email Generator Service: Clean.
appuserimplemented correctly. - Gateway: Clean.
appuserimplemented correctly. - Scheduler: Implements
appuser. Missingscrapyinrequirements.txt, meaning it cannot trigger subprocesses. - Scraper Service: Perfectly implements
appuserwith explicitPLAYWRIGHT_BROWSERS_PATHcache adjustments to securely share Playwright binaries.
- Are all services present? Yes.
- Are healthchecks defined and correct? Yes, standard
curlandpg_isreadyoperations. - Are depends_on relationships correct and complete? Missing
postgresrequirement forcrawler(but crawler actually doesn't use postgres, only redis). Gateway appropriately depends on everything. - Are environment variables wired correctly between services? Mostly yes. Missing
redis_datavolume map. - Are volumes defined for persistent data (Postgres, Redis)? Postgres was mapped to
pgdata, but Redis data persistence wasn't mounted anywhere locally. - Are ports correctly mapped and documented? Yes.
- Will docker-compose up actually start the full system successfully in its current state? It will start, but the scheduler will crash on Scrapy subprocess executions, and Contact Discovery Playwright scraping will error out due to binary permission issues.
| Service | Unit tests | Integration tests | Contract tests | E2E tests | Overall |
|---|---|---|---|---|---|
| crawler | β partial | β missing | β missing | β missing | 10% |
| scraper | β partial | β missing | β complete | β missing | 30% |
| aggregator | β complete | β complete | β complete | β missing | 80% |
| contact | β missing | β missing | β missing | β missing | 0% |
| email-gen | β partial | β missing | β missing | β missing | 20% |
| gateway | β missing | β missing | β missing | β missing | 0% |
| scheduler | β missing | β missing | β missing | β missing | 0% |
Five most critical missing tests:
- Contact Discovery E2E: Requires network mock handling to simulate GitHub and Clearbit endpoints. This system dictates pipeline conversion.
- Gateway E2E Routing tests: Given the gateway proxies the entire stack, unit tests testing httpx orchestration rules are paramount.
- Contact Discovery Data insertion: Validating deduplication constraint handling in Postgres across edge cases.
- Email Generation LLM fallback: Assuring Ollama API disconnection safely yields to Jinja2 backup templates.
- Scheduler Process mocking: Assuring APScheduler instances successfully boot and cycle without dying silently on arbitrary subprocess failures.
- Severity: CRITICAL
- File and line:
scheduler/tasks.py:118 - Description: Scheduler runs
subprocess.run(["scrapy", "crawl", spider]). Scrapy is not installed in the scheduler's Docker container. - Suggested fix: Eject Scrapy from the Scheduler entirely. Hit an HTTP endpoint on a Crawler container daemon, or use Docker's API to spin up ephemeral crawler containers.
- Severity: HIGH
- File and line:
contact-discovery-service/Dockerfile:28 - Description:
playwright installruns beforeUSER appuser. The binaries land in/root/.cache/ms-playwrightwhichappusercannot read. - Suggested fix: Mimic the approach in
scraper-service/Dockerfilewhich setsENV PLAYWRIGHT_BROWSERS_PATH=/home/appuser/.cache/ms-playwrightand runschown.
- Severity: HIGH
- File and line:
gateway/main.py:231 - Description:
await asyncio.sleep(3)assumes background contact discovery is complete within 3 seconds. It almost certainly won't be on the first run for a given company due to rate limits. - Suggested fix: Replace arbitrary sleep with a polling loop, or utilize WebSockets for real-time pushing.
- Severity: MEDIUM
- File and line:
contact-discovery-service/verifier.py:39 - Description: Uses an in-memory dictionary to track rate limiting per domain (
_domain_rate). If Docker replicas scale horizontally, rate limits will desync. - Suggested fix: Migrate rate-limiting logic to Redis
INCRandEXPIRE.
- Severity: LOW
- File and line:
aggregator-service/main.py:105 - Description: The
GET /jobsendpoint simply returns a list of dictionaries up to the arbitrarylimit. No metadata is provided regarding total database records. - Suggested fix: Wrap API response in a JSON envelope providing
total_countandnext_pagecursors.
- Rewrite
contact-discovery-service/DockerfilePlaywright installation rules. - Fix Scheduler Scrapy crash by altering orchestration parameters.
- Implement basic network authorization via API keys on the Gateway routing configuration.
- Add pagination to the
/jobsand/contactsendpoints. - Rewrite the arbitrary
asyncio.sleepworkflow into a pub-sub UI event. - Replace the
contact-discoverydictionary rate limiter with Redis bounds.
Establish the ai_assist Claude orchestration layer into active pipelines. Introduce UI visualization arrays displaying conversion metrics over historically tracked schedules. Scale scraper implementations bypassing traditional authwalls using residential proxy rotations.
- Add
total_countpagination wrappers to/api/jobs(aggregator-service/main.py). - Add explicit file type upload restrictions mapping magic bytes (
email-generator-service/main.py). - Correct
crawler-service/Dockerfileto employ a non-rootappuser. - Migrate
_domain_rateinverifier.pyto useredis-pystandard implementations. - Abstract generic Jinja2 logic out of
/api/digestinto parameterized payloads.
- Abstract Playwright context closures inside try/finally blocks preventing memory leaks (
contact-discovery-service/discovery.py). - Implement specific Exception catchers replacing global Exception nets across scraper pipelines (
scraper-service/main.py). - Connect the Unstop scraper metadata fully into Gateway proxy channels (
scraper-service/scrapers/unstop.py). - Generate Pytest fixtures modeling Github API responses protecting CI environments (
tests/unit/). - Rewrite Scheduler daemon scraping into direct HTTP POST requests invoking isolated worker routines (
scheduler/tasks.py).
- Rewrite the composite orchestrator
POST /api/workflow/applydiscarding arbitrary sleep timers. - Integrate a basic OAuth or API Key authentication intercept middleware over the Gateway proxy map.
- Expand asynchronous Playwright handling logic combating generic Cloudflare authwalls dynamically without requiring human intervention.
| Variable | Service | Required | Default | Description |
|---|---|---|---|---|
DATABASE_URL |
Aggregator, Contact, Email | Yes | postgresql://... |
Connection URI |
REDIS_HOST |
Aggregator, Crawler, Scraper | Yes | redis |
Message broker host |
REDIS_PORT |
Aggregator, Crawler, Scraper | No | 6379 |
Message broker port |
GITHUB_TOKEN |
Contact | No | None |
Elevates GH API limits |
GMAIL_ADDRESS |
Yes | None |
Dispatch host | |
GMAIL_APP_PASSWORD |
Yes | None |
SMTP validation | |
YOUR_NAME |
No | Applicant |
Profile injection | |
OLLAMA_BASE_URL |
No | http://host.docker.internal:11434 |
Local LLM host | |
CRAWL_INTERVAL_HOURS |
Scheduler | No | 8 |
Operational loop interval |
SCRAPER_WAIT_SECS |
Scheduler | No | 60 |
Post-scrape normalization buffer |
GATEWAY_PORT |
Gateway | No | 8080 |
Public dashboard assignment |
- Clearbit Autocomplete API: Public endpoint. No Auth. Free tier. Used for domain inference.
- GitHub API: Public
/searchand/orgsmapping. Elevated by optionalGITHUB_TOKEN. Limits: 60/hr public, 5000/hr auth. - LinkedIn Public Directory: Playwright rendered. Bounded heavily by bot-protection walls.
- Ollama: Localhost API. Used for drafting templates. Zero external cost.
- Gmail SMTP: Authorized endpoints securely tracking external text submissions.
Dependencies are cleanly segregated into standard configurations inside requirements.txt.
- Outdated: No severely outdated distributions noted.
FastAPIruns efficiently. - Unused: None explicitly noted.
- Note: Run
pip install pip-audit && pip-auditon each service container independently to capture true production vulnerability metrics.
- Findings: The core
docker-compose.ymlmounts sensitive strings directly inside environment configuration blocks (POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:-jobpass}). - Exposure:
.env.examplecorrectly masks literal values. - Hardcoded: None found across the python modules directly. The services cleanly inherit environment tokens.
- SQLi:
asyncpgoperates strictly via parameterized arguments ($1, $2). The query compiler strictly avoids.format()or string concatenation injection vectors. Raw queries are secure. - Path Traversal:
email-generator-service/main.py:280handles arbitrary file uploads. While the file isn't written to disk physically, it relies purely on.endswith('.pdf')rather than checking magic bytes. - XSS: Scraped data relies on the frontend
dashboard.htmlfor sanitization. Depending on JS configuration, un-sanitized job titles could manipulate DOM structures.
- Status: Non-existent.
- Exposed Endpoints: The Gateway (
:8080) actively permitsDELETE /contacts,PATCH /jobs, andPOST /emails/{id}/sendwith zero authentication bounds. Anyone possessing the IP space can delete data or trigger outbound spam campaigns from the user's Gmail address. - Recommendation: Deploying BasicAuth or API Header verification across the Gateway proxy is mandatory prior to cloud hosting.
- Scan commands directly via:
pip install pip-audit && pip-audit - Packages like
FastAPIandhttpxmaintain minimal surface area but should be systematically tracked in CI environments.
- Sandboxing: Playwright actively initializes utilizing
--no-sandboxto operate inside Docker arrays. This implies browser escape vulnerabilities fundamentally jeopardize the entire container runtime. - SSL: Did not observe arbitrary
VERIFY_SSL = Falseconditions, meaning internal proxy MITM architectures remain mitigated.
- Connection:
smtplib.SMTP_SSLsecurely wraps outbound data packets. - Spoofing: Outbound headers map precisely to the initialized
GMAIL_ADDRESS. Header injection vectors remain low due to Python standard libraryEmailMessageabstractions sanitizing parameter data.
- Root Operations:
crawler-serviceutilizes the default root operator traversing its internal operations. - Exposed Ports:
docker-compose.ymlmaliciously exposes5433(Postgres) and6379(Redis) broadly to the0.0.0.0host rather than strictly containing routing to the isolatedjobnetbridge layer.
- Retention: Discovered PII (Names, Emails) strictly maintain permanence in Postgres with zero purging lifecycle attached.
- SBERT Cache: Resumes are parsed instantly but not explicitly stored or encrypted at rest on the database.
| # | Finding | Severity | File/Location | Fix Effort |
|---|---|---|---|---|
| 1 | Lack of Auth on Gateway | CRITICAL | gateway/main.py |
1 hour |
| 2 | Exposed Postgres/Redis host ports | HIGH | docker-compose.yml |
10 mins |
| 3 | Crawler executing as root | HIGH | crawler-service/Dockerfile |
10 mins |
| 4 | File upload bypass potential | MEDIUM | email-generator-service/main.py |
2 hours |
| 5 | Unpurged PII databases | LOW | contact-discovery-service |
1 hour |
Do immediately (before exposing this on any public network)
- Strip
- "5433:5432"and- "6379:6379"fromdocker-compose.yml. - Append
Dependsauthentication routines mapping a generic API Key ontogateway/main.py.
Do before v1.0
- Implement a unified
appuserenforcement standard universally across Docker architectures. - Abstract File uploads to validate true magic byte assignments.
Industry standard hardening (longer term)
# Run SAST and secret validation checks automatically
bandit -r .
trivy image jobhunter_scraper
semgrep --config=auto
detect-secrets scanAdd OAuth logic mapping distinct dashboard control configurations safely over standard JWT bearer paradigms.