Showing posts with label data harvesting. Show all posts
Showing posts with label data harvesting. Show all posts

Tuesday, September 1, 2026

Own a gun? Go to church? Do yoga? AI can find out in seconds.; Politico, September 1, 2026

 ALFRED NG, Politico; Own a gun? Go to church? Do yoga? AI can find out in seconds.

A series of demos on the Hill has both Democrats and Republicans alarmed about artificial intelligence’s ability to plumb commercial databases for information on Americans.

"Lawmakers and staffers from dozens of congressional offices have seen AI’s privacy-busting prowess in demos arranged in recent months by CivAI, a nonprofit that says its aim is to educate the public about AI’s dangers and capabilities.

Besides building troves of personal information gleaned from commercial databases, the dossiers offer advice on potential ways to blackmail, coerce or stalk the individual in question, according to four Republican and eight Democratic staffers who were granted anonymity because they were not authorized to speak on the record.

The demos started landing amid the debate over whether to renew a set of government spying powers that expired in June — legislation that offers the most realistic chance for addressing the issue during this Congress, a Democratic House aide who witnessed one of the briefings said."

Tuesday, July 23, 2024

The Data That Powers A.I. Is Disappearing Fast; The New York Times, July 19, 2024

Kevin Roose , The New York Times; The Data That Powers A.I. Is Disappearing Fast

"For years, the people building powerful artificial intelligence systems have used enormous troves of text, images and videos pulled from the internet to train their models.

Now, that data is drying up.

Over the past year, many of the most important web sources used for training A.I. models have restricted the use of their data, according to a study published this week by the Data Provenance Initiative, an M.I.T.-led research group.

The study, which looked at 14,000 web domains that are included in three commonly used A.I. training data sets, discovered an “emerging crisis in consent,” as publishers and online platforms have taken steps to prevent their data from being harvested.

The researchers estimate that in the three data sets — called C4, RefinedWeb and Dolma — 5 percent of all data, and 25 percent of data from the highest-quality sources, has been restricted. Those restrictions are set up through the Robots Exclusion Protocol, a decades-old method for website owners to prevent automated bots from crawling their pages using a file called robots.txt."