Learning to Blame: Localizing Novice Type Errors with Data-Driven Diagnosis
<p>From the Abstract:</p> <blockquote> <p>Localizing type errors is challenging in languages with global type inference, as the type checker must make assumptions about what the programmer intended to do. We introduce Nate, a data-driven approach to error localization based on supervised learning. Nate analyzes a large corpus of training data — pairs of ill-typed programs and their “fixed” versions — to automatically learn a model of where the error is most likely to be found. Given a new ill-typed program, Nate executes the model to generate a list of potential blame assignments ranked by likelihood. We evaluate Nate by comparing its precision to the state of the art on a set of over 5,000 ill-typed OCaml programs drawn from two instances of an introductory programming course. We show that when the top-ranked blame assignment is considered, Nate’s data-driven model is able to correctly predict the exact sub-expression that should be changed 72% of the time, 28 points higher than OCaml and 16 points higher than the state-of-the-art SHErrLoc tool. Furthermore, Nate’s accuracy surpasses 85% when we consider the top two locations and reaches 91% if we consider the top three.</p> </blockquote>
Read the full article at arxiv.org →
Related stories
- Hacker News · 1 points · about 14 hours ago
- Hacker News · 6 points · about 13 hours ago
- Hacker News · 377 points · 16 days ago
- Are you data driven, or are you just busy?posthog.comHacker News · 1 points · 8 days ago
- Ars Technica · 0 points · 7 days ago
- Hacker News · 3 points · 10 days ago