"In its quest to make ChatGPT less toxic, OpenAI used outsourced Kenyan laborers earning less than $2 per hour, a TIME investigation has found.
"The work was vital for OpenAI. ChatGPT’s predecessor, GPT-3, had already shown an impressive ability to string sentences together. But it was a difficult sell, as the app was also prone to blurting out violent, sexist and racist remarks. This is because the AI had been trained on hundreds of billions of words scraped from the internet—a vast repository of human language. That huge training dataset was the reason for GPT-3’s impressive linguistic capabilities, but was also perhaps its biggest curse. Since parts of the internet are replete with toxicity and bias, there was no easy way of purging those sections of the training data. Even a team of hundreds of humans would have taken decades to trawl through the enormous dataset manually. It was only by building an additional AI-powered safety mechanism that OpenAI would be able to rein in that harm, producing a chatbot suitable for everyday use....
"OpenAI’s outsourcing partner in Kenya was Sama, a San Francisco-based firm that employs workers in Kenya, Uganda and India to label data for Silicon Valley clients like Google, Meta and Microsoft. Sama markets itself as an `ethical AI' company and claims to have helped lift more than 50,000 people out of poverty....
"The data labelers employed by Sama on behalf of OpenAI were paid a take-home wage of between around $1.32 and $2 per hour depending on seniority and performance. For this story, TIME reviewed hundreds of pages of internal Sama and OpenAI documents, including workers’ payslips, and interviewed four Sama employees who worked on the project. All the employees spoke on condition of anonymity out of concern for their livelihoods.
"[F]or all its glamor, AI often relies on hidden human labor in the Global South that can often be damaging and exploitative."
https://time.com/6247678/openai-chatgpt-kenya-workers/
"The work was vital for OpenAI. ChatGPT’s predecessor, GPT-3, had already shown an impressive ability to string sentences together. But it was a difficult sell, as the app was also prone to blurting out violent, sexist and racist remarks. This is because the AI had been trained on hundreds of billions of words scraped from the internet—a vast repository of human language. That huge training dataset was the reason for GPT-3’s impressive linguistic capabilities, but was also perhaps its biggest curse. Since parts of the internet are replete with toxicity and bias, there was no easy way of purging those sections of the training data. Even a team of hundreds of humans would have taken decades to trawl through the enormous dataset manually. It was only by building an additional AI-powered safety mechanism that OpenAI would be able to rein in that harm, producing a chatbot suitable for everyday use....
"OpenAI’s outsourcing partner in Kenya was Sama, a San Francisco-based firm that employs workers in Kenya, Uganda and India to label data for Silicon Valley clients like Google, Meta and Microsoft. Sama markets itself as an `ethical AI' company and claims to have helped lift more than 50,000 people out of poverty....
"The data labelers employed by Sama on behalf of OpenAI were paid a take-home wage of between around $1.32 and $2 per hour depending on seniority and performance. For this story, TIME reviewed hundreds of pages of internal Sama and OpenAI documents, including workers’ payslips, and interviewed four Sama employees who worked on the project. All the employees spoke on condition of anonymity out of concern for their livelihoods.
"[F]or all its glamor, AI often relies on hidden human labor in the Global South that can often be damaging and exploitative."
https://time.com/6247678/openai-chatgpt-kenya-workers/
no subject
Date: 2023-12-10 04:45 pm (UTC)UGH.
no subject
Date: 2023-12-10 05:11 pm (UTC)1) These companies chose to scrape some of the most racist blogs for their training sets. From a Washington Post article listing the training data for Google's C4, we see that both vdare and stormfront were in the training data for Google's generative AI model.
2) The cool technology is really underpaid labor. There are a lot of things that companies want us to consider to be created by cool tech that are brought to you by underpaid labor. There have been stories in the past of refugee labor camps in machine learning.
no subject
Date: 2023-12-10 05:34 pm (UTC)no subject
Date: 2023-12-10 06:03 pm (UTC)