Google Gemini: Everything you need to know about the new generative AI platform

Kyle Wiggers

April 29, 2024 at 7:27 PM·9 min read

Google's trying to make waves with Gemini, its flagship suite of generative AI models, apps and services.

So what is Gemini? How can you use it? And how does it stack up to the competition?

To make it easier to keep up with the latest Gemini developments, we've put together this handy guide, which we'll keep updated as new Gemini models, features and news about Google’s plans for Gemini are released.

What is Gemini?

Gemini is Google's long-promised, next-gen GenAI model family, developed by Google's AI research labs DeepMind and Google Research. It comes in three flavors:

Gemini Ultra, the most performant Gemini model.
Gemini Pro, a “lite” Gemini model.
Gemini Nano, a smaller "distilled" model that runs on mobile devices like the Pixel 8 Pro.

All Gemini models were trained to be “natively multimodal” -- in other words, able to work with and use more than just words. They were pretrained and fine-tuned on a variety of audio, images and videos, a large set of codebases and text in different languages.

This sets Gemini apart from models such as Google's own LaMDA, which was trained exclusively on text data. LaMDA can't understand or generate anything other than text (e.g., essays, email drafts), but that isn't the case with Gemini models.

What's the difference between the Gemini apps and Gemini models?

Image Credits: Google

Google, proving once again that it lacks a knack for branding, didn't make it clear from the outset that Gemini is separate and distinct from the Gemini apps on the web and mobile (formerly Bard). The Gemini apps are simply an interface through which certain Gemini models can be accessed -- think of it as a client for Google's GenAI.

Incidentally, the Gemini apps and models are also totally independent from Imagen 2, Google's text-to-image model that's available in some of the company's dev tools and environments.

What can Gemini do?

Because the Gemini models are multimodal, they can in theory perform a range of multimodal tasks, from transcribing speech to captioning images and videos to generating artwork. Some of these capabilities have reached the product stage yet (more on that later), and Google's promising all of them -- and more -- at some point in the not-too-distant future.

Of course, it's a bit hard to take the company at its word.

Google seriously underdelivered with the original Bard launch. And more recently it ruffled feathers with a video purporting to show Gemini's capabilities that turned out to have been heavily doctored and was more or less aspirational.

Google’s best Gemini demo was faked

Still, assuming Google is being more or less truthful with its claims, here's what the different tiers of Gemini will be able to do once they reach their full potential:

Gemini Ultra

Google says that Gemini Ultra -- thanks to its multimodality -- can be used to help with things like physics homework, solving problems step-by-step on a worksheet and pointing out possible mistakes in already filled-in answers.

Gemini Ultra can also be applied to tasks such as identifying scientific papers relevant to a particular problem, Google says -- extracting information from those papers and “updating” a chart from one by generating the formulas necessary to re-create the chart with more recent data.

Gemini Ultra technically supports image generation, as alluded to earlier. But that capability hasn't made its way into the productized version of the model yet -- perhaps because the mechanism is more complex than how apps such as ChatGPT generate images. Rather than feed prompts to an image generator (like DALL-E 3, in ChatGPT’s case), Gemini outputs images "natively," without an intermediary step.

Gemini Ultra is available as an API through Vertex AI, Google’s fully managed AI developer platform, and AI Studio, Google's web-based tool for app and platform developers. It also powers the Gemini apps -- but not for free. Access to Gemini Ultra through what Google calls Gemini Advanced requires subscribing to the Google One AI Premium Plan, priced at $20 per month.

The AI Premium Plan also connects Gemini to your wider Google Workspace account — think emails in Gmail, documents in Docs, presentations in Sheets and Google Meet recordings. That’s useful for, say, summarizing emails or having Gemini capture notes during a video call.

Gemini Pro

Google says that Gemini Pro is an improvement over LaMDA in its reasoning, planning and understanding capabilities.

An independent study by Carnegie Mellon and BerriAI researchers found that the initial version of Gemini Pro was indeed better than OpenAI's GPT-3.5 at handling longer and more complex reasoning chains. But the study also found that, like all large language models, this version of Gemini Pro particularly struggled with mathematics problems involving several digits, and users found examples of bad reasoning and obvious mistakes.

Early impressions of Google’s Gemini aren’t great

Google promised remedies, though -- and the first arrived in the form of Gemini 1.5 Pro.

Designed to be a drop-in replacement, Gemini 1.5 Pro is improved in a number of areas compared with its predecessor, perhaps most significantly in the amount of data that it can process. Gemini 1.5 Pro can take in ~700,000 words, or ~30,000 lines of code — 35x the amount Gemini 1.0 Pro can handle. And -- the model being multimodal -- it’s not limited to text. Gemini 1.5 Pro can analyze up to 11 hours of audio or an hour of video in a variety of different languages, albeit slowly (e.g., searching for a scene in a one-hour video takes 30 seconds to a minute of processing).

Gemini 1.5 Pro entered public preview on Vertex AI in April.

An additional endpoint, Gemini Pro Vision, can process text and imagery -- including photos and video -- and output text along the lines of OpenAI’s GPT-4 with Vision model.

Using Gemini Pro in Vertex AI. Image Credits: Gemini

Within Vertex AI, developers can customize Gemini Pro to specific contexts and use cases using a fine-tuning or "grounding" process. Gemini Pro can also be connected to external, third-party APIs to perform particular actions.

Google brings Gemini Pro to Vertex AI

In AI Studio, there's workflows for creating structured chat prompts using Gemini Pro. Developers have access to both Gemini Pro and the Gemini Pro Vision endpoints, and they can adjust the model temperature to control the output’s creative range and provide examples to give tone and style instructions -- and also tune the safety settings.

Gemini Nano

Gemini Nano is a much smaller version of the Gemini Pro and Ultra models, and it's efficient enough to run directly on (some) phones instead of sending the task to a server somewhere. So far, it powers a couple of features on the Pixel 8 Pro, Pixel 8 and Samsung Galaxy S24, including Summarize in Recorder and Smart Reply in Gboard.

The Recorder app, which lets users push a button to record and transcribe audio, includes a Gemini-powered summary of your recorded conversations, interviews, presentations and other snippets. Users get these summaries even if they don’t have a signal or Wi-Fi connection available -- and in a nod to privacy, no data leaves their phone in the process.

Gemini Nano is also in Gboard, Google’s keyboard app. There, it powers a feature called Smart Reply, which helps to suggest the next thing you’ll want to say when having a conversation in a messaging app. The feature initially only works with WhatsApp but will come to more apps over time, Google says.

And in the Google Messages app on supported devices, Nano enables Magic Compose, which can craft messages in styles like "excited," "formal" and "lyrical."

Is Gemini better than OpenAI's GPT-4?

Google has several times touted Gemini's superiority on benchmarks, claiming that Gemini Ultra exceeds current state-of-the-art results on "30 of the 32 widely used academic benchmarks used in large language model research and development." The company says that Gemini 1.5 Pro, meanwhile, is more capable at tasks like summarizing content, brainstorming and writing than Gemini Ultra in some scenarios; presumably this will change with the release of the next Ultra model.

But leaving aside the question of whether benchmarks really indicate a better model, the scores Google points to appear to be only marginally better than OpenAI's corresponding models. And -- as mentioned earlier -- some early impressions haven't been great, with users and academics pointing out that the older version of Gemini Pro tends to get basic facts wrong, struggles with translations and gives poor coding suggestions.

How much does Gemini cost?

Gemini 1.5 Pro is free to use in the Gemini apps and, for now, AI Studio and Vertex AI.

Once Gemini 1.5 Pro exits preview in Vertex, however, the model will cost $0.0025 per character while output will cost $0.00005 per character. Vertex customers pay per 1,000 characters (about 140 to 250 words) and, in the case of models like Gemini Pro Vision, per image ($0.0025).

Let's assume a 500-word article contains 2,000 characters. Summarizing that article with Gemini 1.5 Pro would cost $5. Meanwhile, generating an article of a similar length would cost $0.1.

Ultra pricing has yet to be announced.

Where can you try Gemini?

Gemini Pro

The easiest place to experience Gemini Pro is in the Gemini apps. Pro and Ultra are answering queries in a range of languages.

Gemini Pro and Ultra are also accessible in preview in Vertex AI via an API. The API is free to use "within limits" for the time being and supports certain regions, including Europe, as well as features like chat functionality and filtering.

Elsewhere, Gemini Pro and Ultra can be found in AI Studio. Using the service, developers can iterate prompts and Gemini-based chatbots and then get API keys to use them in their apps -- or export the code to a more fully featured IDE.

Code Assist (formerly Duet AI for Developers), Google's suite of AI-powered assistance tools for code completion and generation, is using Gemini models. Developers can perform “large-scale” changes across codebases, for example updating cross-file dependencies and reviewing large chunks of code.

Google's brought Gemini models to its dev tools for Chrome and Firebase mobile dev platform, and its database creation and management tools. And it's launched new security products underpinned by Gemini, like Gemini in Threat Intelligence, a component of Google's Mandiant cybersecurity platform that can analyze large portions of potentially malicious code and let users perform natural language searches for ongoing threats or indicators of compromise.

Gemini Nano

Gemini Nano is on the Pixel 8 Pro, Pixel 8 and Samsung Galaxy S24 -- and will come to other devices in the future. Developers interested in incorporating the model into their Android apps can sign up for a sneak peek.

Is Gemini coming to the iPhone?

It might! Apple and Google are reportedly in talks to put Gemini to use for a number of features to be included in an upcoming iOS update later this year. Nothing’s definitive, as Apple is also reportedly in talks with OpenAI, and has been working on developing its own GenAI capabilities.

This post was originally published Feb. 16, 2024 and has since been updated to include new information about Gemini and Google's plans for it.

Engadget
Apple has reportedly resumed talks with OpenAI to build a chatbot for the iPhone
Apple has resumed talks with OpenAI, the maker of ChatGPT, to build an AI-powered chatbot into the iPhone, according to a new report.
3d ago
TechCrunch
It's a sunny day for Google Cloud
Google Cloud, Google's cloud computing division, had a blockbuster fiscal quarter, blowing past analysts' expectations and sending Google parent company Alphabet's stock soaring 13%+ in after-hours trading. Google Cloud revenue jumped 28% to $9.57 billion in Q1 2024, bolstered by the demand for generative AI tools that rely on cloud infrastructure, services and apps. Google Cloud's operating income grew nearly 5x to $900 million, up from $191 million.
4d ago
TechCrunch
Snowflake releases a flagship generative AI model of its own
All-around, highly generalizable generative AI models were the name of the game once, and they arguably still are. Case in point: Snowflake, the cloud computing company, today unveiled Arctic LLM, a generative AI model that's described as "enterprise-grade." Available under an Apache 2.0 license, Arctic LLM is optimized for "enterprise workloads," including generating database code, Snowflake says, and is free for research and commercial use.
6d ago
TechCrunch
Too many models
Other large language models like LLaMa or OLMo -- though they technically share a basic architecture -- don't actually fill the same role. There's some deliberate confusion about these two things, because the models' developers want to borrow a little of the fanfare associated with major AI platform releases, like your GPT-4V or Gemini Ultra.
10d ago
TechCrunch
Google goes all in on generative AI at Google Cloud Next
This week in Las Vegas, 30,000 folks came together to hear the latest and greatest from Google Cloud. What they heard was all generative AI, all the time. Google Cloud is first and foremost a cloud infrastructure and platform vendor.
17d ago
TechCrunch
Google Cloud Next 2024: Everything announced so far
Google’s Cloud Next 2024 event takes place in Las Vegas through Thursday, and that means lots of new cloud-focused news on everything from Gemini, Google’s AI-powered chatbot, to AI to devops and security. Last year's event was the first in-person Cloud Next since 2019, and Google took to the stage to show off its ongoing dedication to AI with its Duet AI for Gmail and many other debuts, including expansion of generative AI to its security product line and other enterprise-focused updates and debuts. Don’t have time to watch the full archive of Google's keynote event?
19d ago
TechCrunch
Adobe's working on generative video, too
Adobe says it's building an AI model to generate video. Offered as an answer of sorts to OpenAI's Sora, Google's Imagen 2 and models from the growing number of startups in the nascent generative AI video space, Adobe's model -- a part of the company's expanding Firefly family of generative AI products -- will make its way into Premiere Pro, Adobe's flagship video editing suite, sometime later this year, Adobe says. Like many generative AI video tools today, Adobe's model creates footage from scratch (either a prompt or reference images) -- and it powers three new features in Premiere Pro: object addition, object removal and generative extend.
15d ago
Engadget
Google's new AI video generator is more HR than Hollywood
Google Vids is not a replacement for AI-powered video generation tools like OpenAI's Sora. Instead, it uses stock footage and personal documents and media to spit out videos that office workers would be proud of.
21d ago
Engadget
Google Gemini chatbots are coming to a customer service interaction near you
At the ongoing Google Cloud Next conference in Las Vegas, the company has revealed the Gemini-powered chatbots its partners are working on, some of which you could end up interacting with.
21d ago
TechCrunch
Humane’s $699 Ai Pin is now available
Humane today announced the availability of its first product, the Ai Pin. The Bay Area-based hardware startup has been kicking around since 2017, a year after co-founders Bethany Bongiorno and Imran Chaudhri left Apple. Ai Pin is the first of what Humane hopes will be a long line of devices aimed at harnessing the power and popularity of generative AI platforms such as OpenAI’s ChatGPT and Google’s Gemini.
19d ago
TechCrunch
Google launches Code Assist, its latest challenger to GitHub's Copilot
At its Cloud Next conference, Google on Tuesday unveiled Gemini Code Assist, its enterprise-focused AI code completion and assistance tool. If this sounds familiar, that's likely because Google previously offered a similar service under the now-defunct Duet AI branding. Code Assist is both a rebrand of the older service as well as a major update.
21d ago
TechCrunch
Vana plans to let users rent out their Reddit data to train AI
In the generative AI boom, data is the new oil. From Big Tech firms to startups, AI makers are licensing e-books, images, videos, audio and more from data brokers, all in the pursuit of training up more capable (and more legally defensible) AI-powered products. Shutterstock has deals with Meta, Google, Amazon and Apple to supply millions of images for model training, while OpenAI has signed agreements with several news organizations to train its models on news archives.
17d ago
TechCrunch
Google releases Imagen 2, a video clip generator
Google doesn't have the best track record when it comes to image-generating AI. In February, the image generator built into Gemini, Google's AI-powered chatbot, was found to be randomly injecting gender and racial diversity into prompts about people, resulting in images of racially diverse Nazis, among other offensive inaccuracies. Google pulled the generator, vowing to improve it and eventually re-release it.
21d ago
TechCrunch
UK's antitrust enforcer sounds the alarm over Big Tech's grip on GenAI
The U.K.'s competition watchdog, Competition and Markets Authority (CMA), has sounded a warning over Big Tech's entrenching grip on the advanced AI market, with CEO Sarah Cardell expressing "real concerns" over how the sector is developing. In an Update Paper on foundational AI models published Thursday, the CMA cautioned over increasing interconnection and concentration between developers in the cutting-edge tech sector responsible for the boom in generative AI tools. The CMA's paper points to the recurring presence of Google, Amazon, Microsoft, Meta and Apple (aka GAMMA) across the AI value chain: compute, data, model development, partnerships, release and distribution platforms.
19d ago
Engadget
The US and UK are teaming up to test the safety of AI models
The UK and the US governments have signed a Memorandum of Understanding in order to create a common approach for independent evaluation on the safety of generative AI models.
a month ago
Engadget
OpenAI will train its AI models on the Financial Times' journalism
Generative AI is only as good as the training data used to train the models that power it, so AI companies have increasingly been striking deals with news publishers.
13h ago
TechCrunch
Google lays off staff from Flutter, Dart and Python teams weeks before its developer conference
Ahead of Google's annual I/O developer conference in May, the tech giant has laid off staff across key teams like Flutter, Dart, Python and others, according to reports from affected employees shared on social media. Google confirmed the layoffs to TechCrunch, but not the specific teams, roles or how many people were let go. "As we’ve said, we’re responsibly investing in our company's biggest priorities and the significant opportunities ahead," said Google spokesperson Alex García-Kummert.
12h ago
Yahoo Finance
AMD to report Q1 earnings Tuesday as Wall Street looks for jump in AI and PC sales
AMD will report its Q1 earnings after the bell on Tuesday as Wall Street looks for signs of AI and PC sales growth.
11h ago
Yahoo Sports
NBA playoffs: Celtics cruise to Game 4 win over Heat, lose Kristaps Porzingis to calf injury
The Celtics have a closeout game at home, but don't yet know Porzingis' status for Game 5 and beyond.
5h ago
Autoblog
Pour one out for the Ram 1500 Classic in Canada
Ram hasn't said when or if it will discontinue the truck in the U.S., but this move doesn't look good for the previous-generation Classic.
13h ago

News

Life

Entertainment

Finance

Sports

New on Yahoo

Google Gemini: Everything you need to know about the new generative AI platform

What is Gemini?

What's the difference between the Gemini apps and Gemini models?

What can Gemini do?

Gemini Ultra

Gemini Pro

Gemini Nano

Is Gemini better than OpenAI's GPT-4?

How much does Gemini cost?

Where can you try Gemini?

Gemini Pro

Gemini Nano

Is Gemini coming to the iPhone?

Recommended Stories

Apple has reportedly resumed talks with OpenAI to build a chatbot for the iPhone

It's a sunny day for Google Cloud

Snowflake releases a flagship generative AI model of its own

Too many models

Google goes all in on generative AI at Google Cloud Next

Google Cloud Next 2024: Everything announced so far

Adobe's working on generative video, too

Google's new AI video generator is more HR than Hollywood

Google Gemini chatbots are coming to a customer service interaction near you

Humane’s $699 Ai Pin is now available

Google launches Code Assist, its latest challenger to GitHub's Copilot

Vana plans to let users rent out their Reddit data to train AI

Google releases Imagen 2, a video clip generator

UK's antitrust enforcer sounds the alarm over Big Tech's grip on GenAI

The US and UK are teaming up to test the safety of AI models

OpenAI will train its AI models on the Financial Times' journalism

Google lays off staff from Flutter, Dart and Python teams weeks before its developer conference

AMD to report Q1 earnings Tuesday as Wall Street looks for jump in AI and PC sales

NBA playoffs: Celtics cruise to Game 4 win over Heat, lose Kristaps Porzingis to calf injury

Pour one out for the Ram 1500 Classic in Canada