He made $1,000,000 in 30 days brokering data to AI labs
from Build With AI
35m 37s
Ryan Locke and his co-founder launched Polychairs to bridge the gap between private companies and frontier AI labs by securing, de-identifying, and packaging business data for training. Their business model emphasizes depth over volume—prioritizing full workflow data like Slack messages, emails, and CRM entries that reveal employee operations and decision-making. While the average deal is $300,000, high-value transactions can exceed $1 million, with rare cases like the $10 million Spirit Airlines deal. The data market is growing due to rising demand from enterprises—not just AI labs—training their own models, which increases buyer demand and expands the market from a few large buyers to hundreds of thousands. Key success factors include remote-first, white-collar companies with over 50 employees and three years of consistent software use. De-identification is handled by vetted third-party vendors, with on-prem options for high-risk industries, and contracts include exclusivity periods of 24 months to protect buyer value. Brokers can earn a 6% finders fee, and value-additive guidance—like helping companies structure and present their data—is crucial. The industry is evolving with new data types, such as blue-collar headcam footage, and growing regulatory awareness. A key insight is that data licensing is a strategic asset, and its value is tied to exclusivity, context, and future business exits. This ecosystem is expected to grow significantly as AI adoption and enterprise AI training accelerate, creating long-term opportunities for skilled brokers and data specialists.
Ryan Locke, so there's a lot of hype online right now about people making money,
brokering data or selling data to the AI labs.
As I understand, you're someone who's been in the space for quite a while and you've made a lot of money doing it.
So, first of all, why don't you tell us who you are and what is your track record in this space in particular?
>> Yeah, so Polychairs got started around June or July when my co-founder and I were going through the
data licensing process on our own businesses and then really just took off.
The space is growing really quickly and so in our first month, we've tracked a million dollars in revenue and it's just been really growing from there.
>> I'm saying, so a million in your first month and we'll get into the actual business model that you guys are following.
I think there's a lot of questions that I have personally.
I think there's a lot that the audience wants to know about how this whole world works.
But give us maybe one or two takeaways that if somebody listens to the whole thing, what are they actually going to learn?
>> What they're going to learn is a lot of what the data industry is looking for from representatives that are looking to make money.
There's a lot more than just making intros to be really good in this industry and to make the most money.
And then I'd also just say, where are the industries going and some key things that I might know that aren't really public knowledge that could be of use.
>> I'm super curious about this because we see companies like yours and the micro ones of the world just seemingly come out of nowhere and do.
I mean, you guys had a million or seven figures in revenue in your first month.
I mean, that's a huge feat even for somebody that has a track record like you.
So why don't we start by, can you just give us the elevator pitch of this business model?
Like how does it actually work?
>> Yeah, so these frontier AI labs, they need data to improve the models.
They've used all of the data on the internet.
They've bought as many rare books as they can buy.
And so they need private company data to continue their improvements.
And so what we do is we help with securing the rights to those data sets through acquiring them from business owners.
We do the de-identification work to remove any personal identifying information and really just get the package of data training ready.
So that the frontier labs can come by from us and get it immediately on to work with it for their researchers and save their researchers valuable time.
And so we're really like a procurement partner for those labs on it.
>> Now, how much money are we talking here per deal?
Because I've seen some crazy numbers being thrown around.
I've even heard of people who have gotten pitched these crazy numbers and they're like, look, this seems too good to be true.
Like what, you're just going to hand me 300K or 500K or a million dollars just for my company emails or Slack.
I mean, that almost seems too good to be true.
So I guess the first question is, what is the average deal size that you guys are seeing on your end?
>> Average would be around 300,000.
We've done seven figures in a single transaction.
So that's definitely out there.
And then, yeah, I mean, of course, the Spirit Airlines deal went for $10 million to Google.
And that was for a company that's not even in business anymore.
So there's a lot of upper market potential for this as well.
>> Now, can you kind of run me through all the, because again,
I think on Twitter right now, again, there's a lot of hype.
People are saying like, hey, just go find a company that's been in business for 20 years.
It has 30 plus employees.
And then you can just license their data and turn around and broker it to an AI lab.
Like obviously, it's not that simple.
So can you give us the actual, like, what are the, almost like the line item of things that you guys do to transact and to actually close a deal?
Because I think it's way more in-depth than people are making it out to be online.
>> Yeah, let's see, I have a good slide for this.
Okay, so just making the intro is only step one.
The second piece of this is the real technical due diligence side of acquiring data.
Just a big stack of files is not what's interesting to the AI labs.
If, you know, you just have, if you had a wound down company,
but you don't have all of the context from their full business operations.
So if you're missing their slack, for example, that's a very important part of the process.
What the AI labs want is to be able to see how every employee operates inside of a company.
And the different types of tasks that they do.
And being able to watch a workflow from maybe a client email all the way through to how an internal team talks about the email fixes some resolution.
What solutions they propose, which ones do they not go with versus which ones that they do?
Sometimes that's even more interesting than the final answer, right?
And then how does that get all the way back to the client?
So that full trajectory is what's interesting in that that can happen across multiple different systems.
And so when we're doing these technical due diligence calls, a lot of what we're trying to understand is how the operations of a business work.
You know, for my own company, I wish I had known this, but we've mainly operated on just cell phone.
Just pick up the phone and call the homeowner, just pick up the phone and call the contractor or call the sales rep.
And a lot of that data wasn't ended like recorded anywhere.
And so the actual workflow that AI could view just looked like, take it created, take it closed.
I say, yeah.
And then my business partner, for example, he had, you know, all of his subcontractors even on Slack.
And, you know, so I couldn't even get subcontractors to send me invoices on time and let alone that.
And so he had full richness of data as everyone remote only was communicating on Slack.
And so that was really interesting for the labs.
And so it seems like they want to see kind of like you said, they want to see the workflow end to end in the more context.
They can get around that workflow, the better.
So it's like, you know, Slack messages alone may only be worth so much or maybe nothing at all.
But if you can get the Slack messages and the CRM entries and the tickets that have been created along the way.
And the emails that have been sent internally along the way, that's kind of the full package to kind of oversimplify.
Like is that fair to say?
Yeah, that's correct.
So one thing I want to say, Ryan, so obviously like there's so much that goes on under the hood.
It's not as simple as people are making it on Twitter where it's like, hey, find a company that wants to sell their data.
Connect them to an AI lab and cache a six figure check.
Like there's so much that you guys are doing around like anonymizing the data, which I want to get more into here shortly.
Packaging it correctly, really even understanding what you guys have access to.
So one thing I want to say for the audience because there seems to be a lot of people that want to get into the space as it's getting more popular.
If people want to like refer leads to you, Ryan, like if somebody's like, hey, I've got a company that's interested in selling their data.
I don't know what I'm doing, but I'd rather just hand the lead to Ryan and let him take it to the finish line.
I think they can get what is it like 6% commission if they use the link that we're going to put in the description and in the show notes.
If you want to just send Ryan your deals, let him close them and then you get 6%.
Like, is that how that works? Is it that simple for them?
Yes. And we encourage sell side brokers to not only take our finders fee, but to really be value additive to the process.
That's what my recent article was about that I wrote on X was just about how to be value additive and help these business owners organize their data.
Similar to a sell side for a business, preparing their financials, presenting them in the correct light as opposed to potentially leaving money on the table.
And so a lot of these business owners just like selling their business don't know how to present their data.
And so it makes it more difficult for us to even correct, like, get to the best valuation we can for them.
So if you could be value additive and we could pay you a finders fee and they can help pay some type of like management or, you know, advisory fee, then that's that's even better.
Awesome. So yeah, we'll put that link in the description for sure.
People just want to like send Ryan their deals, let him close them and then you just get a kick back.
That's, I mean, that to me, you can make a killing doing just that because I think before we recorded, you had said that your min or is it your average deal is 300 K.
So somebody would get basically 18 K on average for sending you guys a lead that ends up closing.
Is that about the average deal size?
Yes. Yeah.
Okay. And another thing that you mentioned that I think makes a ton of sense is almost looking at this industry like like you're a business broker.
Right? Like when you made that comparison to the M and A world, there's, I think there's going to be a whole industry of companies that pop up that represent companies not in the sale of their business,
but in the sale of their data and a really good broker representing a company in the sale of their data is going to do exactly what a good business broker does advise them on here.
Here's how we got to package your data. Here's how we've got to market your data.
Here's how we've got to present you guys two different buyers, not just go with the first buyer that knocks in your door and offers you a six figure check.
So that makes things a lot clearer to me.
Now, okay, so my next question and this is something I've seen, you know, being commonly echoed across the timeline is the runway for this industry.
Like I'm seeing people say, like, oh, this has got three to six months left before it's just, you know, the labs have bought all the data and this is no longer an opportunity.
What are your thoughts on that?
So there's there's a few things I think that the timeline is going to be longer
And the main reason why is due to exclusivity on the data, so usually a company's anchor
buyer is going to require that the data that we buy goes to them exclusively for a period
of time, which means that if you count the number of frontier labs, there's, you know,
need several of every single industry just to feed like a base layer of demand.
So I do think it's going to last longer than just three months, and especially the more
rare niche the datasets are, it'll probably just move from where more common companies
are less interesting, but then still if you have a 1000 employee distribution company,
well that's going to still be interesting, there's not as many of those.
So that I think there's just going to move up market in terms of what's interesting
to the labs, additionally, there's other types of data in this data as a revenue stream
for businesses is going to become, I think, a more exploding idea, Ali from micro one was
talking about it recently, that the next thing could be egocentric data, which is basically
wearing headcams and collecting robotics training video.
And so companies that have blue collar workers can have their, like employees wear these
headcams and get paid hourly to do that.
That is incredibly interesting.
So, okay, so this, you're kind of convincing me that this space has more legs than I initially
anticipated, because again, I think people are, I mean, again, you would know better than
me, but from what I understand, like the prices are already starting to come down as far
as the offer that you might have made to a business in June might be a lot higher than
the same offer that you would make today just because there's obviously more competition.
But you're saying that overall, there's still going to be a demand, the demand is just
going to shift to either more niche specialized data or to this maybe harder to get data,
which is like actual, I mean, call it like blue collar data for a lack of a better term,
where it's like, hey, let's go strap, strap a camera to our line workers or to our technicians
and get this robotic, this data that we can train robots on.
Like is that kind of the gist of what you're saying there?
Yeah, and that's only one more way. There's another theory out there that enterprises are going
to start to need data to do their own model training. Like Fortune 500s are all kind of taking
this idea. There's a few, like the Thomas Rooters, the law firm and financial firm, I think
they're the financial firm, but they are training their own models and that's going to potentially
become more common. And if you're an example, it's like a $20 million business, where you might
need hundreds of millions of dollars worth of transaction data to train your model.
And the only way to do that is to aggregate several of these data sources to train your own
models. So if that happens, then the world of buyers explodes from maybe 10 to hundreds of thousands.
If enterprises start training their own models, which is another theory for the market too.
That is such such an interesting insight because you saying that makes perfect sense. I mean,
I think we're already starting to see like the Fortune 10s or the Fortune 100s or even the
Fortune 500s of the world starting to train their own proprietary models. And that's I mean,
that's a moat for them. That's a huge asset for them. So it makes sense that they would be willing
to pay for these other private data sets to increase the speed at which they can train their
own in-house models, therefore increase the size of their moat. And to your point, like that opens
up the pool of buyers right now. Maybe there's just like 10 main buyers in terms of the labs,
accounting for like 95% of the data spend. But two years from now or five years from now,
there might be thousands of buyers. And it might just be like a landscape, just like the business
buying landscape right now. Yeah, there's big corporations that buy up tons of small businesses,
but there's tens of thousands, if not hundreds of thousands of people actively looking to purchase
small businesses in 2026. And you're saying that the data market might go that direction. Is that
correct? Yeah, yeah. I think everyone's going to look to train their own models because of the
cost savings of open source. And they'll need the data to do it. So.
Dude, this is so insightful. And I have a billion more questions. One thing I'll save for the audience.
So again, I know we're like we're covering a lot of topics here. Ryan sharing a screen. He's got
some slides that he's sharing here about kind of how they operate as a company. We're going to take
essentially everything we discussed in this episode here. And we're going to just condense it
into like an easy to read almost like playbook style summary that we'll just put in the description
and in the show. It's just like a free download. If anybody wants kind of the 80 20 of this conversation
so that they can either go and look into this industry themselves or again, look to go and partner
with somebody like Orion and just be his lead gen or less. So we'll have that resource available
for you guys. Now, okay, so Ryan, another question that I see people have all the time is what kind of
data are the labs looking for? Yeah. In terms of the sources, I guess, that to even be more specific,
like and not to not to hit you with too many questions at once, but like is Slack more valuable
than teams? Are we looking for email primarily? Like what actual data are we? Are they buying?
That's yeah, as far as that, it's going to be the full corpuses. There might be some labs that work
on a specialty project where they might need 10,000 PDF files for something. But for the most part,
we're going to need the full context data. And so this is something I wrote about in the article
as well, where yes, a company might have 20 years of history, but if within the last three,
they've changed all of their softwares over and only have partial contexts for anything past the
last three years. We can only really score the last three years for our for our valuation
for the full context. So that's super important. Got it. Okay, and I see one of the kind of the
footers on your slide here, depth of history beats raw volume. So that leads me to the question is
what is like your perfect, what's the perfect fit company in your mind? What criteria does that
company meet? Yeah, I would almost say like remote first is an advantage in this world.
You know, I think that it's not to say that this is operationally correct because we got a lot
done by being in person in my company just by even using whiteboards to help make things just happen,
right? But it's really for the data world is it's not even about what happened. It's how it gets
recorded and documented the whole chain of a workflow. And so probably remote first attributes,
white collar would definitely be more interesting because a lot of the work happens in emails,
in Slack messages, on, you know, word files and things like that. So those are probably
the two big attribute points is remote first white collar. Got it. Remote first white collar
are there, is there like a minimum number of employees? Is there a minimum amount of time
that they need to be in business? Like, could you get a little more granular with the criteria?
Yeah. So the market keeps moving up as far as what is more likely to transact. And so 50
employees is kind of the, if you go above that number, you're probably going to be able to get a
deal. Anything below that becomes, you know, a little bit more difficult to transact. There's a lot
of fixed costs involved. And so then the smaller deals are very similar to like how many networks.
There's the same amount of legal and due diligence involved. So why not just do bigger deals?
So 50 plus employees, three years in business, ideally with all the same softwares.
We prefer US based. There are opportunities for non-US based, but English is currently required.
You know, all the de-identification softwares and things like that are currently being built out
for English language. And so even the possibility of translating it incorrectly and missing something
that was personally identifying is just too risky. And then industry-wise, it rotates and fluctuates
around. The labs change their demands regularly. And so, you know, a couple of weeks back,
there was a demand for heavy construction businesses. And they wanted, you know, these huge commercial
contractors. And then recently the demand has been video game companies. And couldn't tell you why,
right? They don't necessarily give us all that information, but they just say go chase it.
And super interesting. Yeah, yeah. Okay, so it sounds like the industry, there's almost like a
flavor of the month when it comes to the industry. But what about the like, the super regulated
industry? So like banking, finance, healthcare, obviously there's HIPAA compliance issues. But I,
from what I understand, some companies still look to buy that data. Like, oh yeah. So I guess that's
the first question is how do you guys treat those industries? Like, let's say you've got a healthcare
company that's willing to play ball. How would you approach that? Well, it's, yeah, healthcare is a
little bit easier. There's currently a lot of practices in place because of like regular research
purposes for health care purposes that have been in the
space to de-identify patient trajectories and patient files so that they could be used
to train the next batch of doctors.
So that one's a little bit less concerning just because those already exist as workflows
and processes in these companies.
It definitely gets more hairy when you start talking about high finance and those other
professional services where confidentiality is typically a play and non-disclosures
where yes, just because something isn't personally identifying, it's potentially still
materially valuable and difficult to even scope out for PII.
And so those high finance legal type of deals, we work with the third party to do on-prem
de-identification.
It's definitely more expensive to do that, but just for the risk of what could happen,
the only reason you should do this is if you feel like it's a safe transaction and that's
something that we add to be able to do that.
Okay, so when you say on-prem de-identification, that's, that was one of my main questions
too and something that I've heard people asking is like, "Well, I heard a guy last week
say, you know, okay, if you're going to license my data, how do I know I'm not going to end
up on the headlines of CNBC as being the guy that, you know, let my patient data or my
customer data not be de-identified properly?"
So like, I'm sure there's so much that goes into it, but like, how do you guys actually,
like, what does the process look like for de-identifying data?
Yeah, so first we use third parties and there's two routes you can either transfer the data
to them to run on their servers and de-identify, which is slightly cheaper but still expensive
or the on-prem is the most expensive option.
And they're using, you know, series of algorithms to look for any personal identifying information,
any nickname associations, any addresses, locations, emails, phone numbers, and they're replacing
them.
So anywhere that it would say "cory," for example, would become employee 632.
And that way the relational context of the data is, like, retained and that's, that's
really important, that relational context piece because some people ask if they can do the
de-identification themselves and it's a hard no because you can destroy the full value
of the data set if you properly de-identify it.
So it's a very complex piece of the pie right now.
And so it's important to do right, obviously to protect the safety, but then also to protect
the value of the data set.
And then who's liable, like if let's say the third party company makes a mistake and
we forget to, you know, anonymize Sally's name and then her health information leaks
out and, you know, it ends up in somebody's clawed like, like, proper results.
I don't know if that can actually happen, but you get the idea.
Yeah.
And that case, like, who's liable?
Is it you guys?
Is it the data anonymizing company?
Is it the company who sold the data?
And I'm sure that this is covered in the contract, but like from a general perspective,
who's liable?
It would, based on the contract, everyone in the chain would be liable.
Yeah.
And, you know, there's some protections for the business owner.
We have been working with an insurance provider to ensure these transactions.
And so we've got to stop.
Yeah.
We underwrote a specific policy for this and so there's a pilot we have going on to test
this out with them, but just based on the contract, you know, it's most likely everyone.
So I feel like that's an insurance product that is going to explode.
Like there's going to be, I'm sure they're probably already are, but just, and the reason
this comes up for me is like, my dad's been an insurance agent for 30 some odd years.
My cousin and uncle have been in the insurance industry for 14 years.
I worked in an agency for, you know, six or seven months at one point.
So that in my mind is like, dang, I kind of want to just go up and start up a specialized
carrier that just, just sells these essentially like data, brokering insurance policies.
If, you know, for lack of a better term, I mean that, that in my mind, he's going to blow
up in the insurance world.
It's just, but it's like, how do you underwrite it?
I guess there's not enough transaction data at this point to accurately underwrite.
Yep.
And so that's what the, we're currently doing like a very small pilot.
So it's not like a full product yet that anybody can get.
So we're testing it.
But then if it does come to be a, you know, a product, I think it's also really interesting
for insurance companies because then now you're, you know, it's basically free lead flow
as well for them, right?
So when the full policies, right, right, right, right, this is so interesting.
Okay.
So what are the, what are the most common objections you get?
You've got a CEO on the hook.
He's got 100 employees, you're, you're dangling a million dollar payout in front of him.
He obviously wants the quote unquote free money for the asset that he already has from virtually
zero effort on his part, but he also wants to make sure that he's not putting the company
or the shareholders or the board in a spot that is, he's going to later regret.
So like what are all the questions him or the CFO or the CTO are nailing you with when
you're in the later negotiation routes?
A lot of it's going to be just some of those contract terms.
What are their protections in the contract timelines?
We haven't really had a no, typically that's kind of self selecting throughout the funnel,
like if they weren't interested in potentially transacting then.
So we haven't had anyone be a no, most of it comes down just the negotiation of terms
in price, they're at the end.
Yeah, the big one I would say is just the de-identification and it's kind of a real big separation of what
we do differently is this third party, it's how we win a lot of deals over some of our competitors.
You know, when you talk about them receiving your raw files, they have even less incentive
to delete those because there's other valuable data sources in there that's potentially
valuable.
Not to say that anyone's doing this but I'm saying it's tough to know what they're doing
once you transfer them so, yeah, that would be, those questions around de-identification
definitely come up and that's what we've optimized for.
And so you typically answer that saying essentially walking them through your process and being
like, look, we use vetted third parties, there is stipulations in the contract that protect
you to an extent should anything go wrong in that process, like, and that's usually enough
for them, they're like, okay, yeah, it sounds good.
Yeah.
Yeah, the, and the third party, the benefit of that is they're only going to have the raw
files for called a week for the time it takes to de-identify it and then they get it, you
know, they get deleted and they get it and so that protects the business owner to know that
before I receive the files, it's all been de-identified.
And so it protects me and the company owner, so it's, I'd say that's, that's the really
good one.
Yeah.
Got it.
Okay.
Now, are there any questions that I haven't asked that you, that you get asked all the time,
like anything I just blatantly missed?
Because I could probably sit here for another two hours and just pick your brain on this,
but I want to make sure I always try to be a stand-in for the audience and say, like, you
know, what are, what are people going to ask, whether they are a company looking to sell
their data, whether there's somebody who sees this as a business opportunity to be a
middleman, anything we're missing here?
The exclusivity periods are probably one of the more talked about things and so some
people are asking for perpetual licenses to these data sets, exclusivity, which I think
that's probably a bit rich for the buyer side to be able to get that.
And that potentially carries weight in an M&A transaction.
So that would be one thing I'd recommend any sales side broker who helps establish for
their clients is, like, hey, we're not going to sign perpetually.
And then moving then forward, what is the term?
Is it 12 months, 24 months, typically it's going to be a 24-month ask from the buyer and
that protects their ability to then turn the data set around and make the money that
they're supposed to, because they just need to protect their ability to turn this into
a product and then sell it out of the chain to the frontier labs.
And if somebody else gets it before the frontier lab, before they sell it to the frontier lab,
then the frontiers are going to be like, hey, I already have this.
So I would say exclusivity terms, typical is probably 24 months, yeah, that would be the
other big piece.
And obviously, the longer you're willing to be exclusive, the more the buyer is going
to be able to pay, but then there's a trade off at a certain point, right?
Like, to your point, I would never sign a perpetual deal.
Like, that in my mind is probably the worst move you could ever make as a seller, because
you don't know what's going to happen down the lot.
You don't know if that might be more valuable at a certain time or another question I have
is like, what if I sold my data or licensed my data and then I want to sell my business?
Like is that going to affect the value of my business that I want to go and physically
sell my entire business?
Obviously that's something I'm going to have to disclose during due diligence.
I mean, how does that work?
Yeah.
It does.
It does come up in due diligence.
And so that's where we're more on the like deal making side, you know, my partner and I's background working through
M&A transactions
We're definitely more of deal makers and understand what the that's a bigger deal for business owners than making a few hundred thousand
so
Like if that's context that we're given from the sales sidebroker that this is something they're looking to do
And in the near future we can time our you know exploration terms to
the exit timeline
And you know help even guide the business owner that's buying the business through what these terms mean
I have another article. I think I wrote it like two or three months ago at this point about this kind of problem about what this is going to look like for M&A
They will definitely start to see this come up more as a due diligence item
because it's like even outside of
This one licensing opportunity. There's probably going to be several others where data is or revenue stream come into play
And that makes sense. I like I like you're framing around like we can use that as a deal point like if we know
This particular owner they want to sell their data today and maybe they don't want to sell the business for another two years
Then we can make sure the exclusivity period is for about two years so that
When it's time to go and sell the business. Yeah, they'll disclose during due diligence
But yeah, we licensed our data two years ago, but that exclusivity period is up. So hey Mr. new buyer
You're buying the business, but you're getting this asset that you can then turn around and sell again with it
So like technically they're getting a nice kicker if that if you if you structure it right is that kind of what I'm hearing
Yeah, yeah, and then I I don't know that I was in just understanding on the financial side
You're probably not going to be able to add it to the EBITDA value right right
Oh look we doubled our we've from 300k a profit to 600k this year because it's a one because it's like a it's a one time right like these are not recurring transactions as far as I understand
Yeah, yeah, great. No, that's the structure
Got it. Okay. No, that makes perfect sense dude. This has been super super eye-opening and I'm I'm glad that I was able to hear it from someone like you who is
In this space because I have a feeling with this I mean kind of being like the wild wild west right now
Or at least it feels like I'm sure you you probably know that better than anyone
There's going to be a lot of people that pop up
Talking about this thing that don't really know what they're doing
So good to hear it from you who's someone doing this at a high level and actually having success with it
So just a couple reminders for the audience. So again, if you want to bring Ryan leads like one
I think one opportunity right now that people have is if you're good at Legion like let's say you're really good at cold email
Or you're really good at content or you're really good at paid ads
Be a Legion shop that just funnels leads to people like Ryan and take a 6% commission on every one and if the average deal size is
300 grand
That's an 18k commission you're getting and I know it's not going to cost you 18k to get a qualified lead in the door
So that's kind of where my mind goes always thinking about how to monetize this opportunity
So again, we'll put links in the description links in the show notes to
Like one click take you straight to Ryan's site and you can start sending him leads
We're also going to put that playbook that I mentioned in the description and in the show notes
If you want just like a nice summary of the 80/20 of everything we covered that'll all be there
Ryan any closing thoughts from you?
I think this is going to be a great space and there's going to be a lot of room for great people to be a part of it and
I want those people to be with poly shares and I'm going to be here to help support them as
They try to make money in this industry. I love it man
Well again, you seem like you know your stuff. I can tell you definitely do
Anywhere else you want to send people it is whether to follow you on acts follow you anywhere else
follow on twitter is great
Perfect and we'll put Ryan's information in the show notes and in the description
So Ryan, thank you so much for your time and for everybody that tuned in if you enjoyed this
Like the video leave a five stop. I mean this was one of the most interesting ones that I've ever done
And this is like 300 plus episodes at this point. So appreciate your time Ryan and I'm sure we'll be talking soon
Awesome
Ryan Locke so first of all who are you and why should people listen to you? I know you're
Yeah deep in the data brokering trenches if you even want to use that language
You're you know you're buying and selling data on behalf of the big AI labs and making a ton of money doing it
So like you know, why should people listen to you? What is somebody going to take away?
Yeah, so first I don't mind the data brokering language
We operate as like the layer that helps aggregate the data
Turn it into training packages that then sell it onto the frontier lab
So we are doing a value lift in the process, but I mean, I think that in every industry
That's of this size of deal
There are brokers just helping guide people through the process
And so I think that it's just appropriate, you know as much as the label doesn't
Uh, you know some people like to avoid it, but it's just appropriate. So I think that that's okay
Um, and you know, we've been around prior to the big spirit airlines trackings action
Um, we were going through this on our own assets me and my co-founder back in July and June
Um, and that's when we kind of understood the space and he's from the venture world and so
We've had some easy connections into the frontier button. Excuse me
Can we reshoot? Sorry. Yeah, no worries. Let me um, I'll even stop. Let me stop
Podcast Summary
Key Points:
Ryan Locke and his co-founder founded Polychairs to act as a procurement partner for frontier AI labs, securing private company data through licensing and de-identifying it for training purposes.
The average deal size is around $300,000, with some high-value transactions reaching seven figures, such as the $10 million Spirit Airlines deal with Google.
AI labs prioritize deep, end-to-end workflow data—including Slack, emails, CRM entries, and internal tickets—over raw volume, valuing context like employee decision-making and client interaction.
Data brokers must conduct rigorous technical due diligence, including full context validation and de-identification by third-party vendors, with on-prem de-identification used for high-risk sectors like finance and healthcare.
Exclusivity periods typically last 24 months, and the data market is expected to expand due to rising demand from enterprises training their own AI models, not just labs.
Companies with 50+ employees, remote-first operations, and consistent software usage are ideal, with data from the last three years being most valuable due to software migration.
Brokers can earn a 6% finders fee, and value-additive support—like helping businesses organize and present their data—is critical to securing better valuations.
Emerging trends include egocentric data (e.g., headcam recordings of blue-collar workers) and regulated industries such as healthcare, where existing de-identification practices enable safer data use.
Summary:
Ryan Locke and his co-founder launched Polychairs to bridge the gap between private companies and frontier AI labs by securing, de-identifying, and packaging business data for training. Their business model emphasizes depth over volume—prioritizing full workflow data like Slack messages, emails, and CRM entries that reveal employee operations and decision-making. While the average deal is $300,000, high-value transactions can exceed $1 million, with rare cases like the $10 million Spirit Airlines deal.
The data market is growing due to rising demand from enterprises—not just AI labs—training their own models, which increases buyer demand and expands the market from a few large buyers to hundreds of thousands. Key success factors include remote-first, white-collar companies with over 50 employees and three years of consistent software use. De-identification is handled by vetted third-party vendors, with on-prem options for high-risk industries, and contracts include exclusivity periods of 24 months to protect buyer value.
Brokers can earn a 6% finders fee, and value-additive guidance—like helping companies structure and present their data—is crucial. The industry is evolving with new data types, such as blue-collar headcam footage, and growing regulatory awareness. A key insight is that data licensing is a strategic asset, and its value is tied to exclusivity, context, and future business exits.
This ecosystem is expected to grow significantly as AI adoption and enterprise AI training accelerate, creating long-term opportunities for skilled brokers and data specialists.
FAQs
Polychairs acts as a procurement partner for frontier AI labs by securing data rights from businesses, conducting technical due diligence, de-identifying data, and packaging it into training-ready datasets that AI labs can use immediately.
The average deal size is around $300,000, though some transactions have reached seven figures, with notable examples like a $10 million deal for Spirit Airlines data.
AI labs value full context data that shows end-to-end workflows, including Slack messages, CRM entries, emails, and internal tickets, to understand how employees operate and make decisions across systems.
Ideal companies are remote-first, white-collar, with 50+ employees, at least three years in business, using consistent software, and ideally based in the U.S. with English-language operations.
Data is de-identified by vetted third-party services using algorithms to replace personal identifiers (like names) with anonymized codes. Liability falls on all parties in the chain, and insurance is being tested to mitigate risk.
The typical exclusivity period is 24 months, giving buyers time to develop and commercialize the data, while sellers avoid perpetual licenses that could reduce future business value.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.