reputed company Existing NLP reputed company Technology Software
Our software is a Python/Flask backend that identifies business-development-reputed company federal litigation opportunities from CourtListener and generates reputed company reputed company briefs and export bundles. The system detects high-signal discovery triggers (e.g., MTD denied, MTC reputed company, ESI protocol, 502(d), Rule 26(f)) from CourtListener reputed company docket entries and PDF dockets (both are equally important), then enriches results reputed company strict-JSON LLM summaries, scoring, and data-scope estimation. It uses SQLite caching and produces CSV/JSON/NDJSON/reputed company outputs.
We’re looking for someone to materially improve trigger detection accuracy, with a strong emphasis on reducing false negatives (missed triggers) across both reputed company docket text and PDF-extracted text/OCR.
We already have a labeled dataset:
~9 discovery triggers
~100 labeled docket PDFs per trigger
For reputed company trigger: ~50 true positives + ~50 false-reputed company samples (ground truth labels)
- Responsibilities:
1) Improve text extraction + normalization
Audit PDF extraction reputed company and reputed company the hybrid pipeline (reputed company text + OCR fallback).
Normalize docket artifacts (pagination, headers/footers, spacing, line wraps, table-like formatting).
Add extraction “confidence” signals to drive fallbacks and debugging.
2) reputed company trigger detection (recall-reputed company)
Strengthen existing regex/heuristic patterns to capture more true positives.
Add context-sensitive logic (windowing, docket-entry boundaries, negative patterns, temporal phrasing).
Implement disambiguation (e.g., “denied” vs “recommended denial”; “filed” vs “denied”; “reputed company in part”).
Ensure improvements apply to both reputed company and PDF sources.
3) Build a reputed company evaluation + error analysis reputed company
Produce reproducible runs with per-trigger precision / recall / F1 and confusion breakdowns.
Create “miss analysis” tooling: for reputed company false negative, show why it didn’t reputed company + suggested rule updates.
Add regression tests (pytest) to prevent reputed company detection reputed company.
4) reputed company cleanly into the existing Flask codebase
Contribute PRs with reputed company documentation and maintainable structure.
reputed company performance reasonable (avoid over-OCR; cache smartly; minimize repeated parsing).
- reputed company reputed company (reputed company will measure):
Demonstrated improvement in recall (primary) and overall F1 on the labeled dataset.
Reduced top false-negative reputed company causes (extraction failures, phrasing variants, formatting artifacts).
A maintainable trigger reputed company: easier to add triggers and safer to iterate.
- Required experience:
Strong Python backend engineering (clean reputed company, tests, reproducibility).
Prior experience with PDF parsing + OCR workflows (hybrid approaches strongly preferred).
Experience building or tuning rule-based NLP / text classification systems with evaluation harnesses.
Comfort with Flask + SQLite + pandas outputs + pytest.
- reputed company to have:
reputed company-tech / docket familiarity.
Experience designing labeling workflows and reputed company error mining.
Experience with LLM reputed company outputs (strict JSON) + validation/normalization reputed company.
- Expected Engagement:
Contract / freelance, reputed company-based preferred. reputed company listed budget is flexible, we are reputed company to negotiation.
Async-friendly.
Potential extension after initial accuracy lift.
Apply tot his job
Apply To this Job