Back to the episode map

Article

GPT-4o Launch Record: Model, Voice, Rollout, and Safety

Trace GPT-4o from the May 2024 omni-model launch through staged ChatGPT voice rollout, the Sky dispute, the system card, and the model's current API state.

Aug 4, 20265 min readBy Dalton Anderson

GPT-4o Launch and Model Record

OpenAI announced GPT-4o on May 13, 2024 as an omni model designed to accept combinations of text, audio, image, and video input and produce text, audio, and image output. The model announcement, live demonstrations, ChatGPT rollout, voice product, safety record, and later retirement from the ordinary ChatGPT text-model picker were separate events.

The launch should be remembered as a direction and chronology, not as proof that every demonstrated capability was immediately available to every user.

timeline
    title GPT-4o launch and product layers
    2024-05-13 : GPT-4o announced and demonstrated
    2024 : Text and vision access expanded in ChatGPT
         : Native audio capability rolled out separately
    2024-08 : GPT-4o System Card published
    2026-02 : GPT-4o text model retired from ordinary ChatGPT access
    2026-07 : GPT-4o still documented in the OpenAI API model catalog

What "omni" meant

Earlier conversational products often chained separate speech recognition, text reasoning, and text-to-speech systems. OpenAI framed GPT-4o as a model trained across modalities, reducing the need to pass every interaction through a text-only middle layer.

That design aimed to improve latency and preserve more information from tone, interruption, images, and audio. It also expanded the risk surface. Audio can reveal identity, accent, emotion, health, background conversations, location clues, and information about people who never opened the application.

The launch demonstrations showed low-latency conversation, interruption, visual understanding, translation, expressive speech, and other multimodal behavior. A demonstration establishes observed behavior in a staged case. It does not establish reliability, access, retention, consent, safety, or fitness in another setting.

The E018 outline cannot prove the model path

The recovered episode outline planned a live demo of "OpenAI's conversational AI." It does not preserve the exact app version, account, model, voice mode, session, or exchange.

ChatGPT already had a voice-conversation product before the native GPT-4o audio experience was broadly available. Source-era product notices distinguished existing voice mode from later GPT-4o audio and video rollout.

The public package therefore should not say the E018 demo used the complete native GPT-4o speech-to-speech system. The recording or a verbatim transcript would be needed to establish what happened.

[[What the Recovered E018 Outline Can and Cannot Establish]] preserves that source limit.

The launch model and the ChatGPT product were not the same record

GPT-4o named a model family. ChatGPT was one product using models and product-specific tools, limits, interfaces, safety layers, data controls, and release schedules.

API access was another surface. Microsoft and other organizations could use OpenAI models through separate products and agreements. A Microsoft Copilot experience was not automatically ChatGPT.

This matters when someone says "GPT-4o could do this in May 2024." The answer may refer to a model evaluation, launch demo, API capability, ChatGPT text experience, preview voice mode, or later product.

The claim should identify the surface and date.

The system card arrived later

OpenAI's GPT-4o System Card describes model preparation, evaluations, external red teaming, and selected mitigations. It covers risks that became more salient with native audio, including unauthorized voice generation, speaker identification, sensitive-trait attribution, emotional reliance, audio robustness, and multilingual behavior.

The system card is a vendor-authored safety report. It supplies evidence about the tests OpenAI described. It does not prove that the model is safe for every language, accent, disability, person, setting, or consequential task.

It was also published after the E018 recording. Later evidence should improve the current record without being projected backward into Dalton's source-era experience.

The Sky controversy needs precise wording

E018's outline says the controversy involved Scarlett Johansson's voice and highlights consent concerns. The stronger reconstructed claim was that the voice was used without permission.

OpenAI stated that Sky was voiced by another professional actor, that the casting process began before outreach to Johansson, and that it paused the voice after concerns. That is OpenAI's account. It is not an independent legal finding.

The reviewed evidence does not establish that OpenAI trained Sky on Johansson's voice. The public page should not say it did.

Several questions remain distinct: whether a synthetic voice is natural, whether it resembles a person, whether a performer authorized the source recording, whether a company cloned a voice, whether users understand the identity, and whether the use is deceptive or violates a right.

The FTC's voice-cloning record discusses fraud, authentication, provenance, and detection. It does not resolve the Sky dispute.

Current state in July 2026

Ordinary ChatGPT access to the GPT-4o text model was retired in February 2026. OpenAI said API access was unchanged at that time.

OpenAI's current API model catalog still documents GPT-4o as of this review. The catalog controls the current API record and should be rechecked before publication.

Current ChatGPT Voice should not be described as the unchanged GPT-4o text model that was retired from the picker. Product voice models, underlying model families, interfaces, and names can differ.

The durable page separates four states: the May 2024 model launch, the ChatGPT rollout, the later safety record, and the current product or API availability.

How to use a launch record

A launch record is useful for historical questions. It answers what the vendor announced, demonstrated, and released at the time. It also records later corrections and changes.

It should not be used as a current recommendation. Before using a current voice or multimodal product, verify the exact account, plan, application, model, settings, data controls, retention, training choices, permissions, accessibility, and task.

[[How to Evaluate a Voice Assistant for Privacy and Consent]] provides the people-and-data method. [[How to Compare AI Assistant Capability With Workflow Fit]] provides the task evaluation.

GPT-4o's May 2024 launch changed expectations about how conversational an AI product could feel. The accurate history keeps that significance without turning the demonstration into timeless product truth.

This launch record was developed with AI assistance from the preserved E018 outline, the linked OpenAI system card and API model record, the FTC voice-cloning record, and the Venture Step source ledger. Dalton Anderson remains the author. OpenAI sources describe OpenAI's model, evaluations, and product state. Model, product, safety, voice, consent, accessibility, legal, source, and founder review are required before publication. Publication is not authorized.

Sources

Follow the evidence.

  1. June 2024 Recall updateblogs.windows.com
  2. Current Recall privacy and controlsupport.microsoft.com
  3. Current GPT-4o API documentationdevelopers.openai.com
  4. Manage Recall for Windows clientslearn.microsoft.com
  5. Recall security and privacy architectureblogs.windows.com
  6. GPT-4o system cardcdn.openai.com
  7. Spotify episodeopen.spotify.com
  8. Current Recall use and requirementssupport.microsoft.com
  9. OpenAI API deprecationsdevelopers.openai.com
  10. Introducing Copilot+ PCsblogs.microsoft.com
GPT-4o Launch Record: Model, Voice, Rollout, and Safety