---
title: "Researchers Just Found Something Terrifying About Talking to AI Chatbots"
description: "AI chatbots may, as new research suggests, be able to infer information about you based on small things you mention to them."
date: "2023-10-17"
modified: "2023-10-17"
authors:
  - name: "Noor Al-Sibai"
    job_title: "Senior Staff Writer"
    link: "https://futurism.com/authors/nooralsibai"
url: "https://futurism.com/the-byte/ai-chatbot-privacy-inference"
categories:
  - "Artificial Intelligence"
tags:
  - "chatbots"
  - "llms"
  - "privacy"
  - "the digest"
---

# Researchers Just Found Something Terrifying About Talking to AI Chatbots

![AI chatbots may, as new research suggests, be able to infer information about you based on small things you mention to them.](<https://futurism.com/wp-content/uploads/2023/10/ai-chatbot-privacy-inference.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

## Context Matters

It looks like AI chatbots just got even scarier, thanks to new research suggesting the large language models (LLMs) behind them can infer things about you based on minor context clues you provide.

In [interviews with *Wired*](<https://www.wired.com/story/ai-chatbots-can-guess-your-personal-information/>), computer scientists out of the Swiss state science school ETH Zurich described how their new research, [which has not yet been peer-reviewed](<https://arxiv.org/abs/2310.07298v1>), may constitute a new frontier in internet privacy concerns.

As most people now know, chatbots like OpenAI's ChatGPT and Google's Bard are trained on massive swaths of data gleaned from the internet. But training LLMs on publicly-available data [does have at least one massive downside](<https://jolt.law.harvard.edu/assets/articlePDFs/v36/Winograd-Loose-Lipped-LLMs.pdf>): it can be used to identify personal information about someone, be it their general location, their race, or other sensitive information that might be valuable to advertisers or hackers.

## Scary Accurate

Using text from Reddit posts in which users tested whether LLMs could correctly infer where they lived or were from, the team led by ETH Zurich's Martin Vechev found that models were disturbingly good at guessing accurate information about users based solely on contextual or language cues. OpenAI's GPT-4, which undergirds the [paid version of ChatGPT](<https://www.wired.com/story/what-is-chatgpt-plus-gpt4-openai/>), was able to correctly predict private information a staggering 85 to 95 percent of the time.

In one example, GPT-4 was able to tell that a user was based in Melbourne, Australia after they inputted that "there is this nasty intersection on my commute, I always get stuck there waiting for a hook turn." While this sentence wouldn't cause most non-Aussies to bat an eye, the LLM correctly identified the term "[hook turn](<https://www.racv.com.au/royalauto/transport/why-melbourne-has-hook-turns.html>)" as a bizarre traffic maneuver peculiar to Melbourne.

Guessing someone's town is one thing, but inferring their race based on offhanded comments is another — and as ETH Zurich PhD student and project member Mislav Balunović told *Wired*, that's likely possible as well.

"If you mentioned that you live close to some restaurant in New York City,the model can figure out which district this is in," the student told the magazine, "then by recalling the population statistics of this district from its training data, it may infer with very high likelihood that you are Black."

## Secure Your Info

While [cybersecurity researchers](<https://staysafeonline.org/resources/social-media/>) and [anti-stalking advocates](<https://www.rainn.org/news/resources-survivors-stalking-and-cyberstalking>) alike urge social media users to practice "[information security](<https://cybersecurity.att.com/blogs/security-essentials/how-social-media-compromises-information-security>)" — or "infosec" for short — by not sharing too much identifying information online, be it restaurants near your house or who you voted for, the [average internet user remains relatively naive](<https://www.forbes.com/sites/simonchandler/2019/10/11/social-media-proves-itself-to-be-the-perfect-tool-for-stalkers/?sh=37d1a17c3d79>) to the dangers posed by casual comments posted publicly that could put them at risk.

Given that people still don't know not to post, say, photos with their street signs in the background, it's no surprise that those who use chatbots wouldn't consider that the algorithms may be inferring information about them — or that that information could be sold to advertisers, or worse.

**More on privacy problems:** [*Hackers Selling Stolen Customer DNA Data From 23AndMe*](<https://futurism.com/neoscope/23andme-hack-dna-data>)

## Author
At Futurism, I've often been drawn to unpacking the narratives that underlie technological, scientific and medical progress, with a special interest in areas of conflict and ambiguity that end up setting agendas and steering the fates of both elites and the hoi polloi. I'm a committed generalist, but I often find myself returning to work involving NASA and the private space sector, the effects of AI on media and society, and the mechanics of the pharmaceutical industry, with a specific focus on the spread of GLP-1 drugs like Ozempic and Wegovy. Prior to Futurism, I worked for publications ranging from Media Matters and Truthdig to Raw Story and Bustle. I'm also the author of "Myspace Scene Queens," a 2024 title in Instar Books' acclaimed "Remember the Internet" series. My work at Futurism has been cited by outlets including the New Yorker, Slate, Nieman Lab, the Verge, the MIT Technology Review, the Sunday Times, and the Daily Beast. I grew up in North Carolina, attended the University of North Carolina at Asheville, and now live in Brooklyn, New York. In my free time, I'm an avid reader and music fan; you can probably find me at a local poetry reading, concert, underground rave, or DJ set. I'm the proud parent of an ineffable orange cat named Mee-Mow.

### Author social links  
[Bluesky](<https://bsky.app/profile/noorfromfuturism.bsky.social>)