The Role
The role is centered on protecting email environments from advanced threat vectors—phishing, BEC scams, and AI-driven deepfakes—by building and maintaining models that run directly in production. You will own models operating at immense scale, running complex aggregations over hundreds of millions of records inside live Microsoft 365 and Google Workspace environments. The core technical challenge isn’t just training models; it’s deploying, fine-tuning, and maintaining agentic workflows and SLMs that process massive email volumes with minimal latency and high precision.
About the Product
The platform provides an AI-driven email security ecosystem designed to detect and neutralize inbox threats in real time. Operating at global enterprise scale, the system combines adaptive machine learning, autonomous agent workflows, and distributed threat signals across Microsoft 365 and Google Workspace integrations. It safeguards thousands of organizations globally by processing hundreds of millions of live email events to block sophisticated, emerging attack surfaces.
Technology Stack: The machine learning pipeline relies on Python as the core language, running across distributed cloud infrastructure on AWS. The environment spans supervised learning models, small language models (SLMs), agentic LLM integrations, NLP frameworks, and rule-based fallback engines. Infrastructure is structured to support continuous monitoring, production model deployments, real-time feature aggregation, and automated drift detection at scale.
What You’ll Be Doing
- Engineer large-scale feature extraction pipelines and anomaly detection models targeting advanced phishing tactics
- Deploy, monitor, and maintain production ML models and SLMs operating over high-volume distributed email streams
- Design and fine-tune autonomous agentic integrations and natural language processing (NLP) systems to counter emerging threats
- Implement real-time model evaluation systems to identify performance drift, edge-case regressions, and false positives
- Establish feature aggregation methods capable of processing and querying data sets exceeding hundreds of millions of records
- Architect pragmatic detection mechanisms, applying rule-based or statistical approaches where traditional machine learning is inefficient
What We Expect
Must-have
- 3+ years of production experience operating as a Data Scientist in high-throughput environments
- Strong proficiency in Python applied to machine learning and large-scale data manipulation
- Solid grounding in machine learning fundamentals, with the pragmatic judgment to select non-ML or rule-based solutions when appropriate
- Demonstrated capability in handling ambiguity and adapting to rapidly changing technological requirements
- Strong technical communication skills for collaborative model deployment and cross-functional problem solving
Nice to have
- Hands-on experience with Natural Language Processing (NLP) techniques
- Familiarity with cloud computing architectures, specifically AWS
- Experience overseeing the complete end-to-end ML production lifecycle
Why This Role Is Worth Your Time
- You are building detection models that operate on active, real-world attack vectors—the feedback loop between engineering new defenses and stopping live cyber threats is instantaneous
- The data science mandate covers an unusually wide technical toolkit: from traditional supervised learning and rule-based logic to modern SLMs, NLP, and agentic workflows
- You will work directly with genuine big-data scale, processing hundreds of millions of records in live production rather than isolated sandbox environments