Pilot track

AI in Software Development

Experiments, RCT, construct validity, and benchmark validity—then an adoption memo.

Track progress0/5

1. Orientation

Decide in ten minutes whether a production case deserves a deeper read.

  1. 01

    BitsAI-CR: Automated Code Review via LLM in Practice

    Break down a paper in 30 minutes

2. How the effect is measured

Compare an enterprise experiment, an RCT, and a productivity proxy: what each study measures, and on whom.

  1. 02

    Achieving Productivity Gains with AI-based IDE features: A Journey at Google

    Claims · product experiments · enterprise evidence

    Break down a paper in 30 minutes
  2. 03

    Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity

    RCT · effect size · external validity

    Break down a paper in 30 minutes
  3. 04

    What's DAT? Three Case Studies of Measuring Software Development Productivity at Meta With Diff Authoring Time

    Construct validity of an engineering productivity metric

    Break down a paper in 30 minutes

3. What breaks

Weigh the risk of generated code and the limits of benchmark validity before an adoption decision.

  1. 05

    Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks

    Benchmark validity · functional/security outcomes

    Break down a paper in 30 minutes
CAPSTONE

AI developer productivity: reconcile the evidence

You can open the capstone directly, but the track completes only after every step in it.

Comparison capstone
Every step of the track is required to complete it; the capstone remains directly accessible.