back

by jeffreyrogers·3y ago·view on hn ↗
I don't see why non-invertibility matters. Lots of useful features are non-invertible.

Edit: and if you are dealing with real data sets or producing real datasets for analysis you will often have only approximations to the thing you want to measure. Determining whether your proxy variable is worth including or how to interpret your results in light of it are necessary skills to develop.

1 comments
The feature is bad. The non-invertibility means that you cannot get back the original data that was used to generate the feature, and try to salvage it.
Sure, that makes it less useful. But why is that so bad that the entire dataset should be discarded and not used, even for uses that don't care about that particular part of the original data?
If you want the dataset, scikit even tells you how to get it. If you just want an example dataset, there are better ones. I mean, this seems somewhat like the Lena debacle: why insist on this particular dataset?