Pilot track
AI in Software Development
Experiments, RCT, construct validity, and benchmark validity—then an adoption memo.
Track progress0/5
1. Orientation
Decide in ten minutes whether a production case deserves a deeper read.
- 01Break down a paper in 30 minutes
BitsAI-CR: Automated Code Review via LLM in Practice
2. How the effect is measured
Compare an enterprise experiment, an RCT, and a productivity proxy: what each study measures, and on whom.
- 02Break down a paper in 30 minutes
Achieving Productivity Gains with AI-based IDE features: A Journey at Google
Claims · product experiments · enterprise evidence
- 03Break down a paper in 30 minutes
Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
RCT · effect size · external validity
- 04Break down a paper in 30 minutes
What's DAT? Three Case Studies of Measuring Software Development Productivity at Meta With Diff Authoring Time
Construct validity of an engineering productivity metric
3. What breaks
Weigh the risk of generated code and the limits of benchmark validity before an adoption decision.
- 05Break down a paper in 30 minutes
Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
Benchmark validity · functional/security outcomes
CAPSTONE
AI developer productivity: reconcile the evidence
You can open the capstone directly, but the track completes only after every step in it.
Every step of the track is required to complete it; the capstone remains directly accessible.