The PhD who ships evaluation.
Mohamed A M Elansary, PhD — candidate for Research Scientist - Frontier Benchmarks. Scientific evaluation rigor for frontier benchmarks/datasets and evaluation data, paired with production scaling.
Scientific evaluation
- PhD research in hard forecast evaluation and uncertainty quantification.
- Multimodel comparisons across basins and hydroclimates.
- Imperfect USGS/NOAA/NASA ground truth, QA, provenance, and HPC reproducibility.
Production delivery
- Agentic LLM systems with retrieval and routing.
- Multi-tenant data isolation and production pipelines.
- Founder/CTO delivery from research-shaped problem to working system.
What I would build
A benchmark slice with task taxonomy, held-out provenance-tracked evaluation data, deterministic checks plus a rubric, frontier-model baselines, difficulty-stratified failures, and a data intervention recommendation—designed for academic collaboration and production scaling.
Honest fit boundary
My PhD is environmental/scientific evaluation rather than ML/NLP; my publication record is AMS/dissertation, not NeurIPS/ICML; and the customer/GTM-facing role is a stretch. I bring Fortune 500 and federal technical communication, but do not claim ownership of a named frontier LLM benchmark.
Status and logistics
Compensation/location: $200,000–$350,000 base; Remote US / NYC-SF hybrid.
DFW · Remote US preferred · EAD full/no sponsorship · take-home/work sample preferred.
Official JD: Snorkel AI / Greenhouse