Please, please, please do not use ‘dplyr::ntile()’ in an analysis that will be submitted to a regulatory agency. It’s secretly an ill-conceived SQL tool that doesn’t do what you probably want it to, and it does not…
3 comments
> Unlike other ranking functions, ntile() ignores ties: it will create evenly sized buckets even if the same value of x ends up in different buckets.
Not being able to specify how ties are handled (and choosing, as far as I can tell, a fairly useless definition) fits the bill.
> Not being able to specify how ties are handled (and choosing, as far as I can tell, a fairly useless definition) fits the bill.
Actually, you can completely specify how ties are handled—and should, in any study designed to be repeatable. The docs for dplyr::ntile tell us that:
> To rank by multiple columns at once, supply a data frame.
So, repeatable, completely specified tiebreaking is as easy as adding a tiebreaker column to the dataset, using whatever strategy makes sense for your study. For example, if we wanted random tiebreaking using R's built-in `runif`, all it takes is one extra line of code:
data_table |>
dplyr::mutate(
tie_breaker = runif(dplyr::n()),
ntile_bin = dplyr::ntile(tibble::tibble(partitioning_key, tie_breaker), n = 4))
Almost 100% of the original author's problems could have been avoided by just reading the docs.Read the full thread on Hacker News →
Related stories
- Static Analysis vs. Taint Analysis: Which One Secures Your Python Code?nocomplexity.substack.comHacker News · 2 points · 1 day ago
- Binary Mutation Analysis of Tests Using Reassembleable Disassemblywww-users.cs.umn.eduLobsters · 5 points · over 6 years ago
- Hacker News · 2 points · 5 days ago
- MiMo-v2.6-Pro: Intelligence, Performance and Price Analysisartificialanalysis.aiHacker News · 136 points · 9 days ago
- Hacker News · 2 points · 11 days ago
- Hacker News · 2 points · 11 days ago