A LLM that behaves like a typical Redditor?
What possible use is that?
Air Canada offering a refund of tree fiddy.
You’ll get your refund eventually but first it will try and gaslight you that Air Canada is a woke mind virus before calling you an asshole and then stalking you.
“instead of the $3.50 refund, I’m also authorized to offer you some June 2025 $350 GME calls.”
If it’s trained on the average Reddit reply: $420.69, nice.
I just want to mark the occasion when my previous comment is on 69 points. Noice.
What possible use is that?
I’ve noticed “has this sub gotten more right wing recently?” posts reaching the top post of the day in the last 6 months or so. r/norge and r/unitedkingdom being examples. You can automate bots that change a subreddit’s consensus on certain topics by bot-spamming threads pertaining to those topics, especially in the first hour of a thread going up. I don’t know if that’s happening, or if it has more to do with the Reddit protest that saw mods abdicate their positions last June and new mods being responsible for the change… but it could also be a bit of both.
Do you propose more bots in order to steer the public opinion? That could indeed generate serious money for reddit I suppose!
A redditor bot is a viable example of a forum member bot.
IMO, I don’t think it can drive topics, but it could make things controversial.
Marketing to terminally online people maybe?
Entertaining puns and pointless jokes.
This is what the 3rd party access to API was really all about.
When API access was allowed , all reddit content was effectively free: They needed to ban 3rd party apps so they could sell the accumulated content. I expect using content to train AI also factors into it.
Is it? Because when you build a bot and just scrape Reddit I don’t think you can just use the content to train AI, just like the New York Times. The API change was definitely to sell more ads and get a higher IPO, but I don’t think it was because of AI.
Am I crazy or are you arguing the same point? Scraping is not the same as API access. They closed off the API to everyone for dubious reasons so they can sell that content (both for ads and AI training)… Right??
No you’re not, the post was editted. The original one said it was all because of AI, the entire reason for the API change was to sell to AI companies.
Edit, now I’m in doubt, because if you edit a post that is shown somehow right?
Edit2, just to be clear my point is that Reddit content was never free, before and after the API change. It’s easier to get the content with a decent API, sure. But it was never free, just like the lawsuit the NY Times started.
Reddit is a trove of user built content under the guise of community. What Spez did was to say “thanks for all the free work, suckers!”, put a price sticker on it, and laughed all the way to the bank.
And this is why I’m not active on any Internet community anymore.Nevermind, I guess I just can’t help myself…And this is why I’m not active on any Internet community anymore,
you typed.
Active as in “creating meaningful contributions and contributing to the overall knowledge base”. I still shit post from time to time.
This is going to be a really weird thing to argue, but I just casually read through a bunch of your comments and they seem like meaningful contributions.
Well, I guess I can’t help myself… I’ll shitpost more from now on 😅
^ this comment right here, officer.
Somebody asked chat GPT to appear to be a normal internet user to populate the comments section to manufacture content for normal Internet users to respond to so that they can continue building up their training models.
You couldn’t see the sarcasm because it was set to “hidden”.
And that is another unintended example of why all of my post history was purged before migration.
What are they odds that they kept it in a backup?
Some 4chan users created a backup bot that auto saves every few hours, so if reddit didn’t do it already, 4chan has been doing it for a while. The bot was originally made for 4chan but repurposed for other websites, reddit included.
Yeah, it’s all too late. Shit, PRISM was 2007, so there’s a copy of everything somewhere. Obviously different ends.
Spez like people are even capable of leeching archive.org and still sell the data which was archived for good intentions.
Depends. If they were smart they backed up every content that had a certain number of upvotes and/or a certain number of paragraphs and/or responses. Just to weed out all the 2-3 word comments that no one interacted with. If OP wrote mostly those then Reddit gives a shit about them deleting those.
Welcome to the club.
Don’t cheat yourself just because there are douches that take advantage…
deleted by creator
This is why I don’t blame anyone for editing/deleting their post history on reddit.
Considering how much of Reddit is already bots, I’m sure this will end fantastically.
deleted by creator
The AI:
"IANAL so could you ELI5, so AITA?
THIS."
Ann frankly, I did Nazi that coming.
Holy shit do I hate that comment
I wish spez had a soul so it could leave his body when sexual assault questions eventually yield the phrase “snuggle struggle.”
It’s gonna be trained on everything, even the stuff from 2009, so I’m expecting less of that and more random ‘my fedora chortles intensify’ word salad
It’s funny you say that because there was a ‘hack’ for chatgpt where you could ask it something like how to build a bomb and it would refuse. But when you added TLDR it would do it.
Reddit is all bots, porn, ads and political shit posts. Good luck getting any useful training content out of that.
Maybe that’s the point? Training the AI to produce the blabbering bullshit that’s preferred in social media?
They don’t care if the AI produced is useful, they just want to milk as much money from their content as they can.
The API changes were almost certainly just the groundwork for this and I called it at the time. The ridiculous pricing model for API access is because it’s aimed at the hottest tech companies, not third party app developers.
The enshittification continues because it’s what neoliberalism demands. They’ll sell your content and the data they have about you and still show you ads, because that’s the most profitable. Ethics and product quality don’t even enter into it.
Liberal market gives end users choice. If they don’t choose, they get the consequences.
This is more like people choosing Trump like types and complaining. Alternative exists, choose it.
“The free market can fix it” is just another neoliberal lie, pushed precisely because it doesn’t work. Rather than holding corporations accountable, it blames the population instead.
The reality is that boycotting businesses isn’t always an option and when it is, it’s usually a luxury. Very few products are domestically and/or ethically produced and when they are, they’re extremely expensive, especially for people being fucked out of every cent by their bosses, landlords and utilities.
It’s why the most hated companies in the world continue to bring in record profits.
Regulations are the real answer, which is why neoliberals oppose them.
I really don’t care about people who behave like they are living in North Korea or who wants a North Korean World to live in.
Even Digg people could say “No, F you” to Digg superstar owners. It is just a damn URL to type.
I wish it would die, because honestly some of the porn was great and Lemmy seems to be the one place on the net that doesn’t specifically ban porn, yet has none of it anyway.
I miss bodyswap and part tf captions…
deleted by creator
“Reddit has given access to YOUR conversations and posts to AI companies.”. FTFY
These were created by people, for peoole, and I will ALWAYS disagree that this data is Reddit’s or any other platforms.
Don’t forget your direct messages aren’t end to end encrypted on Reddit, so now AI will be trained on your craziest “private” conversations
There’s one good news. Reddit didn’t want to pay to move all the old DMs to the new chat infrastructure. So they deleted them.
Pretty sure they just didn’t migrate to the new data structure and didn’t actually delete the raw data. They’re effectively deleted for users but not for Reddit.
now AI will be trained on your craziest “private” conversations
I have no idea what horrible thing this will do to an LLM but I’m kind of curious.
Well to be fair, everything you post and comment on Lemmy can be used in the exact same way
Oh no, all the times I sent or received dodo codes from randos so we could trade animal crossing items. Whatever shall I do?
Edit: I’m gonna leave this here for people to use as a resource against Reddit because it may be worth it to do something actionable.
https://thomashunter.name/posts/2023-06-19-how-to-delete-reddit-account-gdpr-ccpa
Well it’s not yours once you post it on some platform, tbf
With reddits severe bot problem, it’ll be like training on unfiltered sewage. Garbage in, garbage out.
Machines training machines? How perverse!
Damn it. I haven’t deleted my account due to how many people I’ve supported and helped, I stopped using it while ago. It seems I’ll have to.
I wouldn’t bother. They’ll just mark all your stuff DELETED=1 and feed it to their AI anyway.
That’s not a bad idea.
One of the original Reddit memes was quite prescient:
Good thing I scrubbed all of my posts and comments that I could. Fuck that site, straight up and down.
Instead of scrubbing, wordbomb them to screw up any AI training
Oh my sweet summer child
deleted by creator
It will get trained on some comment posts.
Let reddit die. Join Lemmy or /kbin. https://join-lemmy.org/ https://kbin.pub/
And what’s to stop instance owners from selling their data?
The eggs are not all in one basket. Less data to sell.
Thanks to federation, the copies of the eggs are. You can’t stop one instance from selling data sourced from federated content until it’s too late.
The only thing stopping them is the fact that anyone who wants the data can just utilize the federation protocol to take any data they want, and there’s not a lot anyone can do about it. You can’t sell something that’s trivial to get for free.
If the question you’re really asking is “what’s stopping content on Lemmy/Mastodon/etc from being used to train an LLM?” the answer is, nothing.
You can’t put a price tag on it. Nothing is stopping anyone from scraping all of the data for free.
mass user exodus to one of the many other identical Instances. Also, data brokers prolly aren’t interested in going after each Instance because no one instance has enough data to make it worthwhile. Yet again, the fediverse proves its resistance to enshitification.
Lmao, if it gets as big as Reddit then it’s worth scraping. It’s not the fediverse making it less worthwhile, just the size.
Yes, it’s not worth running an instance! So let’s all run one! LOL. It’s so worth it. Fuck reddit.
you OK bud?
I wished they had evil lawyers looking after such stuff and sold strictly opt in data to AI corps. Free for FOSS though.
Where’s my cut?
You signed it all away the moment you scrolled down that EULA 😂
Oh, is that what those things do?
Removed by mod
Funny, I thought that is what the unblockable ads were for.
Removed by mod
No it didn’t.
In before poisoning your comments on Reddit turns into the new protest.