Shikib Mehri

Shikib Mehri

Staff Research Scientist at Google DeepMind, working on post-training and agents. Currently leading post-training for Legal AI — including data for SFT and RL, recursive self-improvement, human/SME annotation, and learning from production and post-deployment signal. Previously Director of Research at Contextual AI, managing a team of researchers working on LLMs and agents, and an Applied Scientist at Amazon, where I helped found the Alexa LLM post-training team (which became Amazon AGI). PhD from CMU LTI [thesis] (2022) and BSc from UBC (2018).

At Contextual AI: (1) Reflective Context Learning accepted at COLM 2026 (2) Grounded Language Model reached #1 on the FACTS leaderboard (VentureBeat) (3) LMUnit open-sourced and achieved #1 on RewardBench2 (4) AgentLens launched as a multi-agent evaluation system (5) Post-training recipe featured as a Google Cloud case study

I most enjoy product- and impact-driven research — work that enables next-generation user experiences through strong problem formulation. Key areas include:

During my BSc at UBC I interned at Meta (shipped first subword neural MT), Microsoft, and Amazon (patent), and spent 2 years in bioinformatics at BC Children's Hospital (graph-based genome representation). My PhD at CMU LTI (2018–2022) was on dialog systems, advised by Dr. Maxine Eskenazi, with another Amazon internship (dialoglue, example-driven '21); I was a TA and research mentor throughout both degrees. On Alexa LLM post-training at Amazon, I led early SFT, the RLHF recipe, and data collection, and owned RM/eval. At Contextual AI I went from Member of Technical Staff (Jan 2024) to Technical Lead Manager (May 2024) to Director of Research (Aug 2025), then joined Google DeepMind in May 2026.