Vana plans to let users rent out their Reddit data to train AI

Kyle Wiggers

Updated April 15, 2024 at 1:02 PM·7 min read

In the generative AI boom, data is the new oil. So why shouldn't you be able to sell your own?

From Big Tech firms to startups, AI makers are licensing e-books, images, videos, audio and more from data brokers, all in the pursuit of training up more capable (and more legally defensible) AI-powered products. Shutterstock has deals with Meta, Google, Amazon and Apple to supply millions of images for model training, while OpenAI has signed agreements with several news organizations to train its models on news archives.

In many cases, the individual creators and owners of that data haven't seen a dime of the cash changing hands. A startup called Vana wants to change that.

Anna Kazlauskas and Art Abal, who met in a class at the MIT Media Lab focused on building tech for emerging markets, co-founded Vana in 2021. Prior to Vana, Kazlauskas studied computer science and economics at MIT, eventually leaving to launch a fintech automation startup, Iambiq, out of Y Combinator. Abal, a corporate lawyer by training and education, was an associate at The Cadmus Group, a Boston-based consulting firm, before heading up impact sourcing at data annotation company Appen.

With Vana, Kazlauskas and Abal set out to build a platform that lets users "pool" their data -- including chats, speech recordings and photos -- into datasets that can then be used for generative AI model training. They also want to create more personalized experiences -- for instance, daily motivational voicemail based on your wellness goals, or an art-generating app that understands your style preferences -- by fine-tuning public models on that data.

"Vana’s infrastructure in effect creates a user-owned data treasury," Kazlauskas told TechCrunch. "It does this by allowing users to aggregate their personal data in a non-custodial way … Vana allows users to own AI models and use their data across AI applications."

Here's how Vana pitches its platform and API to developers:

The Vana API connects a user's cross-platform personal data … to allow you to personalize your application. Your app gains instant access to a user's personalized AI model or underlying data, simplifying onboarding and eliminating compute cost concerns. … We think users should be able to bring their personal data from walled gardens, like Instagram, Facebook and Google, to your application, so you can create amazing personalized experiences from the very first time a user interacts with your consumer AI application.

Creating an account with Vana is fairly simple. After confirming your email, you can attach data to a digital avatar (e.g., selfies, a description of yourself and voice recordings) and explore apps built using Vana's platform and datasets. The app selection ranges from ChatGPT-style chatbots and interactive storybooks to a Hinge profile generator.

Image Credits: Vana

Now, why, you might ask -- in this age of increased data privacy awareness and ransomware attacks -- would someone ever volunteer their personal info to an anonymous startup, much less a venture-backed one? (Vana has raised $20 million to date from Paradigm, Polychain Capital and other backers.) Can any profit-driven company really be trusted not to abuse or mishandle any monetizable data it gets its hands on?

Image Credits: Vana

In response to that question, Kazlauskas stressed that the whole point of Vana is for users to "reclaim control over their data," noting that Vana users have the option to self-host their data rather than store it on Vana's servers and control how their data's shared with apps and developers. She also argued that, because Vana makes money by charging users a monthly subscription (starting at $3.99) and levying a "data transaction" fee on devs (e.g., for transferring datasets for AI model training), the company is disincentivized to exploit users and the troves of personal data they bring with them.

"We want to create models owned and governed users who all contribute their data," Kazlauskas said, "and allow users to bring their data and models with them to any application."

Now, while Vana isn't selling users' data to companies for generative AI model training (or so it claims), it wants to allow users to do this themselves if they choose -- starting with their Reddit posts.

This month, Vana launched what it's calling the Reddit Data DAO (Digital Autonomous Organization), a program that pools multiple users' Reddit data (including their karma and post history) and lets them decide together how that combined data is used. After joining with a Reddit account, submitting a request to Reddit for their data and uploading that data to the DAO, users gain the right to vote alongside other members of the DAO on decisions like licensing the combined data to generative AI companies for a shared profit.

We have crunched the numbers and r/datadao is now largest data DAO in history: Phase 1 welcomed 141,000 reddit users with 21,000 full data uploads.
— r/datadao (@rdatadao) April 11, 2024

https://platform.twitter.com/widgets.js

It's an answer of sorts to Reddit's recent moves to commercialize data on its platform.

Reddit previously didn’t gate access to posts and communities for generative AI training purposes. But it reversed course late last year, ahead of its IPO. Since the policy change, Reddit has raked in over $203 million in licensing fees from companies, including Google.

"The broad idea [with the DAO is] to free user data from the major platforms that seek to hoard and monetize it," Kazlauskas said. "This is a first and is part of our push to help people pool their data into user-owned datasets for training AI models."

Unsurprisingly, Reddit -- which isn't working with Vana in any official capacity -- isn't pleased about the DAO.

Reddit banned Vana's subreddit dedicated to discussion about the DAO. And a Reddit spokesperson accused Vana of "exploiting" its data export system, which is designed to comply with data privacy regulations like the GDPR and California Consumer Privacy Act.

"Our data arrangements allow us to put guardrails on such entities, even on public information," the spokesperson told TechCrunch. "Reddit does not share non-public, personal data with commercial enterprises, and when Redditors request an export of their data from us, they receive non-public personal data back from us in accordance with applicable laws. Direct partnerships between Reddit and vetted organizations, with clear terms and accountability, matters, and these partnerships and agreements prevent misuse and abuse of people’s data."

But does Reddit have any real reason to be concerned?

Kazlauskas envisions the DAO growing to the point where it impacts the amount Reddit can charge customers for its data. That's a long ways off, assuming it ever happens; the DAO has just over 141,000 members, a tiny fraction of Reddit's 73-million-strong user base. And some of those members could be bots or duplicate accounts.

Then there's the matter of how to fairly distribute payments that the DAO might receive from data buyers.

Currently, the DAO awards "tokens" -- cryptocurrency -- to users corresponding to their Reddit karma. But karma might not be the best measure of quality contributions to the dataset -- particularly in smaller Reddit communities with fewer opportunities to earn it.

Kazlauskas floats the idea that members of the DAO could choose to share their cross-platform and demographic data, making the DAO potentially more valuable and incentivizing sign-ups. But that would also require users to place even more trust in Vana to treat their sensitive data responsibly.

Personally, I don't see Vana's DAO reaching critical mass. The roadblocks standing in the way are far too many. I do think, however, that it won't be the last grassroots attempt to assert control over the data increasingly being used to train generative AI models.

Startups like Spawning are working on ways to allow creators to impose rules guiding how their data is used for training while vendors like Getty Images, Shutterstock and Adobe continue to experiment with compensation schemes. But no one's cracked the code yet. Can it even be cracked? Given the cutthroat nature of the generative AI industry, it's certainly a tall order. But perhaps someone will find a way -- or policymakers will force one.

Engadget
OpenAI will train its AI models on the Financial Times' journalism
Generative AI is only as good as the training data used to train the models that power it, so AI companies have increasingly been striking deals with news publishers.
3h ago
Engadget
Apple has reportedly resumed talks with OpenAI to build a chatbot for the iPhone
Apple has resumed talks with OpenAI, the maker of ChatGPT, to build an AI-powered chatbot into the iPhone, according to a new report.
3d ago
TechCrunch
Watch it and weep (or smile): Synthesia's AI video avatars now feature emotions
Generative AI has captured the public imagination with a leap into creating elaborate, plausibly real text and imagery out of verbal prompts. Now, Synthesia — one of the ambitious AI startups working in video, specifically custom avatars designed for business users to create promotional, training and other enterprise video content — is releasing an update that it hopes will help it leapfrog over some of the challenges in its particular field. Unlike other generative AI players like OpenAI, which has built a two-pronged strategy — raising huge public awareness with consumer tools like ChatGPT while also building out a B2B offering, with its APIs used by independent developers as well as giant enterprises — Synthesia is leaning into the approach that some other prominent AI startups are taking.
4d ago
TechCrunch
Amazon wants to host companies' custom generative AI models
AWS, Amazon's cloud computing business, wants to become the go-to place companies host and fine-tune their custom generative AI models. Today, AWS announced the launch of Custom Model Import (in preview), a new feature in Bedrock, AWS' enterprise-focused suite of generative AI services. The feature lets organizations import and access their in-house generative AI models as fully managed APIs.
6d ago
TechCrunch
This Week in AI: When 'open source' isn't so open
This week, Meta released the latest in its Llama series of generative AI models: Llama 3 8B and Llama 3 70B. Capable of analyzing and writing text, the models are "open sourced," Meta said -- intended to be a "foundational piece" of systems that developers design with their unique goals in mind. "We believe these are the best open source models of their class, period," Meta wrote in a blog post.
9d ago
TechCrunch
Generative AI is coming for healthcare, and not everyone's thrilled
Generative AI, which can create and analyze images, text, audio, videos and more, is increasingly making its way into healthcare, pushed by both Big Tech firms and startups alike. Google Cloud, Google's cloud services and products division, is collaborating with Highmark Health, a Pittsburgh-based nonprofit healthcare company, on generative AI tools designed to personalize the patient intake experience. Amazon's AWS division says it's working with unnamed customers on a way to use generative AI to analyze medical databases for "social determinants of health."
15d ago
TechCrunch
Adobe's working on generative video, too
Adobe says it's building an AI model to generate video. Offered as an answer of sorts to OpenAI's Sora, Google's Imagen 2 and models from the growing number of startups in the nascent generative AI video space, Adobe's model -- a part of the company's expanding Firefly family of generative AI products -- will make its way into Premiere Pro, Adobe's flagship video editing suite, sometime later this year, Adobe says. Like many generative AI video tools today, Adobe's model creates footage from scratch (either a prompt or reference images) -- and it powers three new features in Premiere Pro: object addition, object removal and generative extend.
14d ago
TechCrunch
Google goes all in on generative AI at Google Cloud Next
This week in Las Vegas, 30,000 folks came together to hear the latest and greatest from Google Cloud. What they heard was all generative AI, all the time. Google Cloud is first and foremost a cloud infrastructure and platform vendor.
16d ago
TechCrunch
Humane’s $699 Ai Pin is now available
Humane today announced the availability of its first product, the Ai Pin. The Bay Area-based hardware startup has been kicking around since 2017, a year after co-founders Bethany Bongiorno and Imran Chaudhri left Apple. Ai Pin is the first of what Humane hopes will be a long line of devices aimed at harnessing the power and popularity of generative AI platforms such as OpenAI’s ChatGPT and Google’s Gemini.
18d ago
TechCrunch
Meta trials its AI chatbot across WhatsApp, Instagram and Messenger in India and Africa
Meta has confirmed to TechCrunch that it is testing Meta AI, its large language model-powered chatbot, with WhatsApp, Instagram and Messenger users in India and parts of Africa. The move signals how Meta plans to tap massive user bases across its various apps to scale its AI offerings. Meta announced plans to build and experiment with chatbots and other AI tools in February 2023.
18d ago

News

Life

Entertainment

Finance

Sports

New on Yahoo

Vana plans to let users rent out their Reddit data to train AI

Recommended Stories

OpenAI will train its AI models on the Financial Times' journalism

Apple has reportedly resumed talks with OpenAI to build a chatbot for the iPhone

Watch it and weep (or smile): Synthesia's AI video avatars now feature emotions

Amazon wants to host companies' custom generative AI models

This Week in AI: When 'open source' isn't so open

Generative AI is coming for healthcare, and not everyone's thrilled

Adobe's working on generative video, too

Google goes all in on generative AI at Google Cloud Next

Humane’s $699 Ai Pin is now available

Meta trials its AI chatbot across WhatsApp, Instagram and Messenger in India and Africa