How much does it bother you that OpenAI is trained on your data? What can we do about it?

duncesplayed · edit-2 3 years ago

How much does it bother you that OpenAI is trained on your data? What can we do about it?

curioushom · 3 years ago

The issue is that most of the content posted is archived fairly quickly. Deleting/rewriting only hurts the humans that might have gone looking for it. The way I look at it is, if the data is searchable/indexable by search engines (as a proxy for all other tools) at any point of its life cycle then it’s essentially permanent.

unfazedbeaver · 3 years ago

That all true. The idea isn’t to remove yourself from the internet. Once you post to the internet, it’s there forever. No, what I am proposing is to hurt reddits chances of being a viable first party resource to train AI.

curioushom · 3 years ago

Unless you’re able to compel a platform to remove your data through something like EU right to be forgotten then the data will remain (in training sets or otherwise). If third parties are able to archive your data, reddit will surely have access to their own archival data and will use the original and edited content for training and let machine learning sort it out.

I’m not saying this to be a defeatist, we need better data ownership and governance laws. Retroactively obfuscating the data will not serve the purpose and provides a false sense of control, which I contend is worse.