---
title: "In Leaked Audio, Microsoft Cherry-Picked Examples to Make Its AI Seem Functional"
description: "A Microsoft researcher giving an internal presentation on an early demo of an AI tool admits that the responses were \"cherry-picked.\""
date: "2024-01-13"
modified: "2024-01-13"
authors:
  - name: "Frank Landymore"
    job_title: "Contributing Writer"
    link: "https://futurism.com/authors/flandymore"
url: "https://futurism.com/the-byte/microsoft-cherrypicked-ai-examples"
categories:
  - "Artificial Intelligence"
tags:
  - "generative ai"
  - "large language models"
  - "microsoft"
  - "the digest"
---

# In Leaked Audio, Microsoft Cherry-Picked Examples to Make Its AI Seem Functional

![A Microsoft researcher giving an internal presentation on an early demo of an AI tool admits that the responses were "cherry-picked."](<https://futurism.com/wp-content/uploads/2024/01/microsoft-cherrypicked-ai-examples-1.jpg>)
*\<em\>Image: Getty / Futurism\</em\>*

## Pick and Choose

Microsoft "cherry-picked" examples of its generative AI's output after it would frequently "hallucinate" incorrect responses, [*Business Insider* reports](<https://www.businessinsider.com/microsoft-cherry-picked-outputs-security-copilot-ai-product-hallucinations-2024-1>).

The scoop comes from leaked audio of an internal presentation on an early version of Microsoft's Security Copilot, a [ChatGPT](<https://futurism.com/the-byte/chatgpt-limit-silly-goose>)-like AI tool designed to help cybersecurity professionals.

According to *BI*, the audio contains a Microsoft researcher discussing the results of "threat hunter" tests in which the AI analyzed a Windows security log for possible malicious activity.

"We had to cherry-pick a little bit to get an example that looked good because it would stray and because it's a stochastic model, it would give us different answers when we asked it the same questions," said Lloyd Greenwald, a Microsoft Security Partner giving the presentation, as quoted by *BI*.

"It wasn't that easy to get good answers," he added.

## Halluci Nation

Functioning like a chatbot — you type a query into a chat window, you get an answer in the style of a customer service rep — Security Copilot is largely built on OpenAI's GPT-4 large language model, which also underpins Microsoft's other generative AI outings like the Bing Search assistant. According to Greenwald, Microsoft got early access to GPT-4, and these demos were "initial explorations" of the tech's capabilities.

Not unlike the [Bing AI](<https://futurism.com/bing-ai-images-women>), which during the early stages of its release would return [responses](<https://futurism.com/microsoft-bing-ai-threatening>) [so](<https://futurism.com/bing-ai-names-enemies>) [insane](<https://futurism.com/bing-ai-sentient>) that it had to be "[lobotomized](<https://futurism.com/microsoft-limited-bing-ai>)," the researchers said that Security Copilot frequently "hallucinated" incorrect responses during its early iterations — a problem seemingly endemic to the technology.

"Hallucination is a big problem with LLMs and there's a lot we do at Microsoft to try to eliminate hallucinations and part of that is grounding it with real data," Greenwald said in the audio, "but this is just taking the model without grounding it with any data."

In other words, the LLM Microsoft used to build Security Copilot, GPT-4, wasn't trained on cybersecurity specific data at the time. Instead, it was used straight out of the box, relying only on its standard — but still immense — general dataset.

## Cherry on Top

Sharing another set of security questions, Greenwald revealed that "this is just what we demoed to the government."

It's unclear if Microsoft used these "cherry-picked" examples in its presentations to the government and other potential customers — or if its researchers were this candid about how the examples were chosen.

A Microsoft spokesperson told *BI* that "the technology discussed at the meeting was exploratory work that predated Security Copilot and was tested on simulations created from public data sets for the model evaluations," adding that "no customer data was used."

**More on AI:** *[Microsoft’s Stuffing Talking Generative AI Into Your Car](<https://futurism.com/the-byte/microsoft-tomtom-generative-ai>)*

## Author
At Futurism, my work has often centered on bringing a sense of clarity and insight to complex topics ranging from the regulation of emerging technologies to the esoteric ideologies of Silicon Valley executives, while striving not to lose the poetic sense of awe inspired by often-obscure fields like astrophysics and quantum computing. I broke the story of CNET using AI to produce articles that turned out to be riddled with factual errors and plagiarism — a dam-breaking inflection point, as I've reported, that's inspired copycats and endless discourse while beguiling stakeholders ranging from tech giants to purveyors of spam around the web. My work at Futurism has been cited by publications including CBS News, the Los Angeles Times, Vice, Gizmodo, Engadget, the Verge, and Vanity Fair. I grew up in locales ranging from India to China, and now live in the exotic suburbs of Virginia. In my free time, I'm an avid reader of weird sci-fi literature, an aficionado of East Asian cinema, and, regrettably, a relapsed gamer. Allegedly, I’m working on a debut novel, currently untitled.

### Author social links  
[Bluesky](<https://bsky.app/profile/f-w-l.bsky.social>)