The End of the Open Internet
Some of the largest headlines in tech recently involves AI and other AI-related topics. Another big story is the closing of the Reddit free API, and in the same area the temporary shutting down of the ability to view tweets while logged out of Twitter. A lot of other tech companies have already taken steps like these to bring up walls, moats and barbed wire to anyone not logged into, and authenticated to their services.
I believe it’s only a matter of time before other similar companies are starting to seriously think over their decision to let logged out viewers access the content on their site. Just imagine when the ability to analyze video and audio the same way text is being analyzed today become a tool for the average Joe. What implications will that have for data heavy streaming services like Youtube, Vimeo, etc.? You might prompt the AI to summarize the best bits from a 2 hour video on Youtube into 10 minutes. Youtube still has to host the 2 hour version and show it to your AI while you only have to spend 10 minutes. This way you can watch several days worth of video in a matter of minutes of hours.
The story so far
To give some backstory to the developments here are a few bullets to consider:
-
Late 2022 - release of ChatGPT 3 to the public
-
Early 2023 - ChatGPT, Midjourney, and many more similar AI and LLM tools emerge on the market and the public quickly catches on.
-
A few weeks later ChatGPT is one of the largest websites in the world with user numbers almost as high as the top social media platforms.
-
The following weeks and months the startup industry catches on and a massive amount of tools hit the market every single day.
-
The open source community does not lag behind but comes up with tools such as GPT4All and AutoGPT. Everyone and their aunt can now have. A local ChatGPT running on a laptop at home. Offline!
-
It. Is. Not. Slowing. Down.
The sources of training data
What is it that all of the above is built upon? Training data.
And where do we get training data from? The internet!
More specifically? Big ass social media sites!
Since the largest websites and social media sites have the most amount of users contributing to the site, they also have the most data sitting on their servers. Due to the nature of their business, the data also have to be accessible to the rest of the world. This is what makes it a huge target to the new AI companies who scrapes the internet in search of raw data to use in their new AI tool. Since they have vast amounts of data and there are an increasing amount of companies that want that data, the servers must be running quite hot at Twitter and Reddit!
Average users get stuck in the middle
The new development in AI has allowed more and more people to collect more and more data since the AI itself will train on the data and come out better, or more specified at the other end. This is what AI training amounts to, collect a lot of data and letting the AI sift through it and keep the information it deems good enough, or unique enough.The social media companies have had enough
This is why Twitter had to temporarily turn off access to anyone not logged in to Twitter, and this is also probably a large part of the reason Reddit is turning off free API access and closing down their services to third party apps, etc. It’s simply a matter of not wanting to pay for thousands of startup’s AI training data. Of course, normal people who just want to browse Reddit on their favourite third party-app will get stuck in the middle, as is often the case.
Want to read more on Tech and AI?