95 points•ingve•4 days ago•47 comments•

47 comments

articulatepang3 days ago
I prefer a slightly more general rule: Make Illegal States Unrepresentable. The "Parse, don't validate" rule is a special case of MISU.

What's the difference? MISU applies even when there's no parsing-like transformation happening. For example, if you have a variable that represents the current state of a network connection, and let's say it can be Disconnected, or Connected to some IP address (this is an oversimplification).

Then one way to do it would be

  struct {
    connected: bool,
    peer_ip: int32
  }
The trouble is that this allows us to represent an illegal/meaningless state: we're disconnected but there's still some junk old peer_ip hanging in there. Even worse, we might have written

  struct {
    connected: bool,
    peer_ip: Option<int32>
  }
Now we could have connected = true but peer_ip = None.

The solution is to use a sum type:

  type connection = 
     Disconnected
   | Connected of int32
(sorry for using made-up syntax; I hope it's clear to anyone familiar with Rust.)

"Make Illegal States Unrepresentable" applies throughout your program, at every interface between modules or functions in the program, including but not limited to parsing input.

laszlokorte3 days ago
In rust this would simply be

   enum ConnectionStste {
        Connected(u32),
        Disconnected
    }
slopinthebag3 days ago
or

   struct Connected { peer_ip: u32 }
   struct Disconnected;
   struct Connection<State> { state: State }
ykonstant3 days ago
Rawwrr, use dependent types and write your entire program in the signature!
Fluorescence4 days ago
Not sure that type is good advice:

    pub struct NonEmpty<T> {
        pub head: T,
        pub tail: Vec<T>,
    }
You'd have to manually implement the traits to support the ergonomics of slices and iteration and costly reallocation if you need to pass ownership as a Vec:

I'd expect:

    pub struct NonEmpty<T> {
        v: Vec<T>,
    }
The constructor would enforce the invariant and then you'd impl Deref and DerefMut for [T] to gain normal len/is_empty/indexing/iteration, passing as &[T] to other funcs and mutating values (which can't break the invariant).

To mutate length while preserving the invariant it's dealers choice e.g.

- add .into_vec() for unwrap/mutate/rewrap

- add invariant preserving mutators of your choice

Rusky3 days ago
There's a whole follow up post about this: https://lexi-lambda.github.io/blog/2020/11/01/names-are-not-...
Fluorescence3 days ago
Not sure switching between languages makes for a compelling argument:

"Look how easy it is to accidentally bypass the invariant of a rust newtype by transliterating the data shape into Haskell and deriving a new type". Uh, ok.

If comparing the "risk of accident" between a newtype wrapper whose only role is enforcing the the invariant versus manually reimplementing vector and iterator semantics to use a different layout... I'd say the newtype wins that.

It would be good advice to keep a newtype that enforces an invariant as a single purpose primitive type. A building block and not a place to add other features.

There might be times I'd prefer structural enforcement e.g. something serialisation related. Converting into a non-rust format is what they are doing in their "accident"!

shim__3 days ago
DerefMut would allow you to call `clear()` on `v` violating the invariant

I'd be great is there were a way to shadow methods but even then guarantees would be poor since Vec might add a new method in the future which isn't covered by invariant checks

Fluorescence3 days ago
clear() is a method of std::Vec not std::slice.

DerefMut to [T] not Vec<T>.

eptcyka4 days ago
Which deref must I use to get most of the existing interface sans `retain()`?
Fluorescence3 days ago
Deref/DerefMut enables implicit type coercion rather than exposing an interface. You can choose the target type and immutable/mutable but not parts of the target type.

You can use all the slice reference methods (that do not require ownership) with:

    impl<T> Deref for NonEmpty<T> {
        type Target = [T];

        fn deref(&self) -> &Self::Target {
            &self.v
        }
    }

    impl<T> DerefMut for NonEmpty<T> {
        fn deref_mut(&mut self) -> &mut [T] {
            &mut self.v
        }
    }
https://doc.rust-lang.org/std/primitive.slice.html

If you DerefMut to a Vec then you won't be able to preserve the invariant.

If you want control over methods to expose then you need wrapper methods for those you want. If you want to expose some of the traits the inner type implements then there are likely derive macros available e.g. with derive_more you could expose just indexing as:

    #[derive(Index, IndexMut)]
    struct MyVec(Vec<i32>);
jelder4 days ago
This is great. Alexis King actually stated that, had she known how popular “Parse, Don’t Validate” had been, she would have written it in a language more widely used than Haskell.
esafak4 days ago
With its rich type system, Haskell is the perfect language to demonstrate the dictum.
bunderbunder4 days ago
With its crap type system Python might be even better, in a strange way.

Haskell's strong, static, non-reflective type system tends to make "parse, don't validate" produce code that also looks nicer. Which is great. So great that it steals a bit of the main message's valor.

In Python, though, it's really easy to just let your data be a dynamically typed list of dicts forever. So easy that parsing into something more strongly typed looks like a whole lot of extra effort. Upon looking at that sort of thing many a working Python programmer, myself included, hears the voice of GvR murmuring disparaging things about "academic" programmers down in the pit of their brain.

Which creates an opportunity to demonstrate all the ways the (arguably) more Pythonic way is actually a royal PITA when you try to make your code robust. Handling and reporting data validity errors gets scattered all over the code, which makes it annoying to maintain. Unit test suites get bloated because it's not obvious what inputs a function should be able to handle. Comments and docstrings to help keep track of this stuff begin to proliferate.

OptionOfT3 days ago
On the footnote:

> [1] Other languages - like Go or Python - have a runtime check that raises some sort of exception or panic when lst[0] is accessed on an empty list or slice.

Rust has the same thing. Accessing a `Vec` by index goes via the index trait: https://doc.rust-lang.org/std/ops/trait.Index.html#tymethod....

Vec implements Index here: https://github.com/rust-lang/rust/blob/d080e7dff1b0fc5454154...

Vec's Index defers to slice's: https://github.com/rust-lang/rust/blob/d080e7dff1b0fc5454154...

Slice defers to... intrinsics: https://github.com/rust-lang/rust/blob/d080e7dff1b0fc5454154...

Which injects a bounds check: https://github.com/rust-lang/rust/blob/d080e7dff1b0fc5454154...

The bounds check: https://github.com/rust-lang/rust/blob/d080e7dff1b0fc5454154...

So I'd say the footnote is not correct.

eliben2 days ago
You are technically correct :) But that's not the point of the footnote, and you're right that I should update it. The point here is that Python and Go don't have the equivalent of "first", and have the programmer rely on the runtime error in lst[0]. Rust can affort to have a "first" method because of Option.
Supermancho3 days ago
"parse don't validate" can be rephrased: "expect the type validation from a parser"

Which is plainly moving the problem around, for types. The value validation is a much simpler problem, as a separate application-specific check.

cpburns20093 days ago
Thank you! I always found the quip "parse, don't validate" to be more confusing than helpful. Parsing doesn't remove the need for validation. All it does is ensures the type is right, but does nothing to validate the values for the domain.
reamaer3 days ago
It is also about clear demarcation, of validated stuff vs not yet validated.

Crystal clear clarity is a nice thing to have. (As with everything, there are trade-offs)

Read the full thread on Hacker News →

Related stories