The Tech and Tradecraft Behind Open Source Intelligence
45m 33s
The Cogs of War podcast episode delves into the evolution of open source intelligence, highlighting the shift towards using specialized tools and platforms for data collection and analysis. The episode features insights from industry leaders like Ryan Curran from XeroFox, who discuss the importance of leveraging software platforms to gather various types of open source data. These platforms enable organizations to gain a holistic perspective on potential threats and prioritize actions based on AI-driven analyses. The discussion emphasizes the challenges of navigating complex information environments, the need for human verification in data analysis, and the technical tradecraft required for successful open source intelligence operations. Additionally, real-world examples illustrate how companies like XeroFox assist public sector organizations in monitoring and addressing threats such as scams and fraudulent activities targeting executives and high-profile individuals.
Transcription
8320 Words, 48857 Characters
You are listening to the Cogs of War podcast, which brings you the best information and analysis
on defense tech and the defense industry.
It's brought to you by Warren The Rocks and our friends at Booz Allen.
My name is Ryan Evans.
I'm the founder of Warren The Rocks.
Open source intelligence has matured from the days when it was simply about gathering
whatever could be scraped from the open web.
Today it depends on specialized tools, platforms, and services that collect, filter, analyze,
and protect the flow of data at a scale no individual analyst can manage.
This episode brings together three leaders from the private sector who deal with these
issues very closely.
We have Ryan Curran, the principal director of private sector services at XeroFox.
Tucker Moore, a vice president at Booz Allen.
And Scott Petrie, co-founder and executive chairman at Authenticate.
Enjoy the conversation.
Starting with you, we've known each other for a few years, obviously.
But how did you first become a technologist and a builder of technology?
Wasn't this company?
It wasn't this company.
No.
I'm a product of Silicon Valley.
Born and raised in the valley.
In the late '80s, I BSed my way into a job at a small local computer manufacturer called
Apple Computer.
I found a path in the networking group where we were building sort of early stage local
area networking and early stages of wide area networking and started out as a network administrator,
ran a competitive analysis lab.
And then over time, became a product manager and started working on sort of shaping products
for markets.
And that led me on my path into other companies and then starting my own companies.
Your last company before this, tell us about that briefly.
I started a company in 1999, which was innovative on a technology front and a business model
front.
I started a company that was SaaS before, it was called SaaS.
We were doing in the cloud before, it was called the cloud email security.
We were a very large processing layer for email where we would do unpacking of messages
and content inspection in real time and make a determination.
Is it good?
Is it bad?
Is it spam?
Is it not?
Does it have confidential information and apply a set of rules to it?
That grew into a pretty big business and then we sold that to Google in 2007.
It was called Postini.
And what made you say, "I've had this incredible exit, incredible business.
It's time to do something quite different," which is rooted in cybersecurity, open source
technology, open source intelligence rather.
It's a great question and the answer is sort of, I didn't decide.
We kind of were presented with an opportunity.
So when we started this company, we thought we would do the same thing we had done to
the SMTP protocol, which is sort of process it in real time as it's live and make a determination
on it.
We started to do similar things around HTTP or the browser-based protocol, so we figured
out how we would be able to sit in the middle of a browser conversation and assess whether
something was good or bad, and we thought we would sell that to enterprise customers.
Growing up in Silicon Valley, you hear the adage, if you build a better mousetrap, etc.
It's not really true.
The CISO market wasn't really ready for what we were building, but we were approached by
an intermediary that said, "We have some government customers that might be interested in this
process for this thing called managed attribution.
You've built a virtualized, scaled infrastructure for processing web-based content in real time.
Could you apply it to this problem of managed attribution?"
Looking for angles for our business, we said, "Sure," and then figured out what that meant
after the fact.
So it was about 2015, 2016, where we really dug in and started shaping our platform to
be a set of resources for analysts to get out on the web, access and process content
without revealing their identity or affiliation.
That's great.
How about you and your company?
How did you get started in technology?
I went to grad school at Hopkins, got my master's, and somehow found myself at a cyber startup
that was looking at network topology and what we now know as a tax service intelligence.
In that model, in the platform, there was a lot of interest in the public sector space
and the way it could be used to look at their own enterprises, their third-party vendor
ecosystems as well.
With that knowledge and that base-level understanding, we also then started adding on technical indicators,
open-source information, and those types of data feeds as well to give a better perspective
of the vulnerabilities, the botnets, and the attackers that are looking at those spaces
and going after and attacking those spaces with malware, botnets, ransomware, and attacks
like that as well.
As we grew our public sector business, I, by default, became the public sector guy and
then you acquire many hats in that growth path, which was a lot of fun.
I ended up being a product manager, program manager, analyst, all those different hats
you wear when you're trying to work with the public sector, especially within a startup
environment.
We really started to shape the platform and the deliverables out of the platform for the
public sector customers because they were one of the early adopters of the technology
and saw the value and what it could bring and their understanding of the attack service
and the attack services they care about, both their own as well as the critical infrastructure
of the nation.
Booz Allen, of course, has pretty prolific capabilities when it comes to intelligence
and open source intelligence, both analysts and investments in technology.
Tell us about how the company approaches that and your role in all that.
As you mentioned, Booz Allen has been supporting the defense intelligence community, federal
law enforcement, specifically from an open source intelligence standpoint for a really
long time.
I think it was in the last probably five years or so where we realized that the convergence
of open source data, cyber and AI has never been more evident.
That Venn diagram has really closed in.
One of the things that we've shifted our focus towards is leveraging those technologies to
make open source collection and exploitation really the key focus.
One of the things that we've done at Booz Allen is really started to shift back from
the vernacular of open source analysis.
Really that's the production and the dissemination piece.
That's the piece that the policymakers or the end customers really care about.
From an end product, what we started really shifting our focus towards is the technology
enablement at the front end, at the middle, and even at the back end from a pre-planning
at a collection standpoint.
There are a lot of considerations that have to be taken into account from a technical
signature perspective.
What does the information environment look like?
Why is that?
Well, the information environment is fast, is becoming more contested, and in some cases
is completely denied.
Help me just to ask a dumb question, but to provoke a thoughtful answer.
Why isn't it just as simple as getting out there on the Internet using your Chrome browser
and finding cool stuff?
That's what some open source often is, or you see people on social media getting on telegram
or getting, pulling stuff off of X and repackaging it and geolocating it, and does it really
need to be that technically complicated?
Absolutely.
It absolutely does.
We joke that open source used to be just Google it, and it has really dived quite a bit deeper
than that because the information environment, the Internet, isn't just bifurcated.
It's trifurcated.
There are multiple different environments.
What an Internet user looks like in one environment is completely different than what a normal
Internet user or a netizen looks like in another environment.
If you really want to be able to traverse that information environment and really get
at the information that is meaningful, you need to know what right looks like.
You need to, A, have access, but, B, you need to know what right looks like, and, C, there's
a technical implication to the tradecraft that you have to play.
You need to be behaving sort of.
Behaviorally, absolutely.
It's almost like if you're a physical spy, you need to act a certain way in a certain
country.
You're saying it's the same way online, too.
You need to make sure that the attributes you're putting off, which are discernible signals
as an Internet user, make you look like just some schmo from St. Petersburg.
Let's just say.
100%.
I mean, there's a difference between obfuscation and hiding in plain sight.
Knowing the difference between those two from a technical standpoint is critical.
I think just as critical is knowing when to use those as well.
When you can simply obfuscate your identity, who you are, and get into those environments
works in a lot of cases.
In some cases, the "hiding in plain sight" and blending in with the millions of other
normal netizens in some of these environments is actually critical as well.
And then being able to collect at scale on those things, because the attack vectors are
so varied and vast now that you have to be able to go do it at scale and actually derive
meaningful information from it.
So it's not just the gathering of it, but also the discerning and analyzing what's in
that data and then determining or prioritizing which things are most important out of that
data set and how you're going to apply those in your environment that's going to have some
kind of meaningful outcome or some determination of reduction of risk for your organization.
And of course, every organization might have a little bit different determination of risk.
Brian, you asked a simple question, I want to just sort of back up and give a really
simple answer.
When you access resources on the Internet, you're known to the person you're connecting
to.
When you're doing that on behalf of an organization with a particular mission, that becomes risky
just from an information leakage perspective.
You don't want to be known.
So you need some level of obfuscation.
Non-attribution, as Brian mentioned, is important, but sometimes you want to look like you're
someone else in a particular region with a particular persona.
Getting access to those sites gets harder as anti-bought mitigation mechanisms get put
in place and people are doing more validation and validating across multiple data signals
across a particular user.
Is there an example you could give?
Reputation of IP, first-time visit without having a "Remember Me" cookie, you get the
cloudflare capture.
Perfect example there.
Or have you tried to sign up for a social media site recently?
The KYC steps you have to go through are ridiculous.
You need to have a cell phone, you need to have an identity, I need to have an alternate
Internet address that you can use reliably to get messages.
The bar for access, even before we get into the countermeasures story, where an adversary
might want to be doing something against you, the bar just to get access to these resources
is increasing.
If you look at that with another trend, which is I don't have math on this, so I'm riffing,
more data used to be on the open web before than now.
Now more data is on what's called the deep web, requires an authentication step to get
in.
Before you can get to Twitter's closing doors, Facebook's closing doors, you need to have
identities on those resources in order to go and access that information.
There are estimates that say it's an order of magnitude more data on the deep web than
on the open web.
That puts a further challenge in front of you.
Not the dark web, but just things you have to log into.
Just things you have to do, an authentication step, whether it's the public library or whether
it's a social media site.
The last thing I'll say very quickly is that when you start thinking about countermeasures,
now you actually, given the nature of these protocols, you actually have an adversary
that's delivering potentially active code back into your environment.
When you get a cookie dropped that allows you to get your stock ticker across your finance
page, there's active code executing on your computer that's fetching data from another
source and that can be an avenue for an exploit into your environment.
It used to be as simple as go Google it and that's open source.
Now not only is it more difficult, but the stakes are much higher from an information
leakage or from a countermeasures perspective.
One of the challenges with that too is as you start to protect against those types of
things that you mentioned, every bit of protection from a traditional cybersecurity standpoint
adds an additional layer of complexity from a technical signature perspective and so those
things in certain places look anomalous.
Those are things that from a technical story perspective, you have to take into account
and so when we talk about two-factor authentication and KYC, even the stories associated with
those from a technical signature perspective are absolutely critical from an end-to-end
collection investigation success story.
And the additional teams and tools you need to go accomplish those tasks at the end of
the day are another thing you have to plan out, process, workflow, budget for, account
for.
It's not as simple as you're just Googling it anymore.
You have to really think through what the process is going to be like.
Not your father's open source intelligence.
Not your father's open source intelligence.
Not collecting newspapers anymore.
Yeah.
I know what you mean because a lot of those reports, that's what you were looking at.
They would just go right and they'd Google and they'd cobble it together and they would
slap maybe an analysis component on top of it and that was what was being published.
And at the time, that provided some value, but it's much more complex now and the process
is you get to go through to access the data to your point and then derive information that's
going to be important to you is a whole other layer of complexity.
Good data is expiring milk, right?
And so from an actionability standpoint, that information environment is changing.
The data is constantly changing and being updated and so being able to get there when
you need to get that data, ex-fill that data securely, and then be able to derive actionable
answers, intelligence, whatever it may be is absolutely critical as well in that timely
manner because otherwise that data either disappears, gets locked behind closed doors
or everything changes.
There's another important distinction here that's emerging, as he said, which is the
time life of the data is short and there's also a proliferation of alerting or scraping
technologies that will synthesize information and inform someone of something, tipping and
queuing as to go look at a particular resource.
This is where our space in the market fits nicely with technology providers like XeroFox,
which is they can do a large-scale access synthesis and help inform analysts or decision
makers as to where threats might emerge.
At some point, provenance needs to be established, whether it's, I need to go verify this or
I need to go capture information in a non-repudiated fashion so I can use it in an evidentiary
proceeding, or if I need to do a little additional phishing, so-and-so send out a signal, it looks
like they're doing something weird.
Let me see where else that person is popping up and I can spider out from that investigation.
The ability for large-scale collection tools and AI analysis and the synthesis that comes
out of that still requires the need for a human to go out there and verify and establish
provenance of what's going on.
Yeah, the age-old speed versus accuracy, there still needs to be human in the loop somewhere.
You need the speed to go collect, but you also need a human in the loop somewhere to
do some level of provenance and make sure the data is accurate so you can make a decision
about what you're going to do with it or inform the decision makers above you who need that
information so they can make a decision about their comfortable risk level.
Well, Scott, you mentioned something really important there about a human in the loop.
In the age of automation and being able to go grab data at scale, being able to process
and exploit it at scale, those are the places where automation, where OSINT is a perfect
landscape for automation, but when we talk about the active collection in many cases,
tradecraft, like good collection tradecraft, and now the technical tradecraft that is required
to do this successfully is incredibly difficult to automate as well, and so making sure that
the people actually understand, have these skill sets and the expertise to know how the
internet works, what the information environment is doing, how they are monitoring your presence
and what they are looking for is incredibly important, along with the traditional, you
know, human-like tradecraft to go in and actually grab something and pull it back out.
So let's talk a little bit about the tech stack of the ideal OSINTR and starting with
where your company's -- we talked about this earlier in the founding, what they actually
offer to that technical stack, starting with you, Ryan, always a bit of a risk to have
another Ryan on the show, but we went ahead and did it anyway.
We'll make it work.
So XeroFox has a software-based platform that gathers open source, deep dark web, social
media from a variety of different sources, a variety of different types of indicators,
actor TTPs, malware, PII, the Compromise Account Credentials, things like that.
So what that really allows is sort of to give a broader, more holistic perspective about
the various types of information that's out there that could impact your organization.
But then again, obviously that's also very important to understand what your actual environment
looks like, both your owned assets as well as your assets that are out there in the cloud,
vendor, the ecosystem, the critical infrastructure you're a part of.
And so the mapping together of the collection and the analysis and then the overlay on top
of your attack surface is really what helps derive the actual information that you need
to go be aware of.
And then the sort of the AI component on top of that is helping you prioritize and score
and prioritize which things are more important than others for you to consider to go be actioned.
And obviously there's a variety of different ways you can do that, but the AI and the platforms
basically allows you to very quickly understand what the type of threat is, where it came
from, how it was derived, how it was collected, and score it in a manner which helps you prioritize
it in your environment and how you're going to action it.
Can you give me, if not like a real-world example, what a real-world example would look like?
Like let's say an analyst has a problem, what is that problem?
What are they trying to solve for?
How are they using your solution?
Sure.
So we have a variety of different customers in the DOD, IC, and federal civilian space
on the public sector side.
And a lot of our public sector customers are concerned about what information is out there
about their executives or their brand.
So there's a lot of different scams, Roman scams, fraud scams, actor selling playbooks
on how to scam the government, things like that.
And then also what information is out there about an executive or a high-risk profile individual?
And they leverage the platform in our analysts to go do deep dives, travel advisories, person
of interest investigations, as well as give a holistic perspective on what information
is out there about you as the executive so that you and your protection teams can protect
those executives from both a cyber and a physical perspective and be more aware of like the
proliferation of that digital persona out there on the internet and how people are trying
to attack it or leverage it to be exploited and ultimately scam citizens of the U.S.
Are you able to discern through your solution who these people are actually doing it and
find out real identities behind some of these accounts that are involved in these scams
or attempts at fraud?
In some places, in some sources, we gather the information, store a catalog, analyze
it, prioritize it.
And then we also have humans in the loop who can do deeper dive investigations to make
some of those more connective tissues.
I mean, again, there's like billions of indicators back there in the back end, right?
So the analysis of the platform does that at scale, at speed to deliver the alerts you
might need to know about.
And then if you need a deeper dive, there's analysts in the loop as well that can give
you that deeper dive to connect those personas, what this actual actor or this screen name
is, and who that person is in real life.
Interesting.
And Scott, where does Authenticate's offering sit in?
How does it work?
So I'll start with a description of what sort of the platform does, and then I'll spend
a minute on a real-world example.
So we start where Ryan said something needs to be actionable.
Our world is really today defined by somebody needing to go out and find something out.
That can be in response to an RFI or a request for information.
Here's a ticket.
Go figure this out.
Go answer this question.
Or it could be an investigator who wants to go out and explore and do their sort of capillary
searches.
The first step is you need to be able to get access to the web and internet-connected applications
like social media, Telegram or Signal or whatever.
And you need to do that in a way that allows you to not reveal who you are, what your affiliation
is, et cetera.
You need to be able to do that oftentimes in particular regions as well.
You want to pop out in France looking like a French-speaking person with a machine that
looks like it lives in France as opposed to somewhere in the Northeast corner of the US.
You need to be able to do that securely.
These web content protocols are active.
They will deliver code that executes on your machine.
That code is designed by some arbitrary alternate party.
So you need to use somebody else's computer to execute these searches and be able to interact
with them at arms length.
You don't want to do it from your own computer.
You need to have a set of applications that allows you to process or synthesize the information.
In the French example, I don't speak French, so I need translation for French things.
And if I grab a video, I need to transcribe it in English so that I could look at it.
I still need applications though that allow me to preserve that original information in
a non-repudiated fashion.
So if I need to use it in an evidentiary proceeding, I can present that.
I can put that in my case file.
So you need a set of tools above and beyond this managed attribution and secure access.
Now you need a suite of tools that work together to allow you to collect and process your information.
And importantly, and more and more importantly in today's world with increasing sophistication
of policies around things like stumbling across U.S. persons data for government employees,
you need to stay within policy guidelines.
You need to have barriers around what a user is authorized to do.
Those need to be defined based on their mission.
You get Eastern Europe, you get Asia Pacific, and your two tools don't cross-connect.
It's a simple example.
And you need to have everything that the analyst does be audited.
Okay.
So how does that work?
What does that work?
Let's say I'm working on financial fraud, and I have a couple of principles that I'm
looking at.
I need to go out and figure out how are they conducting some of these illicit transactions?
What do I want to do?
I want to get on some of the cryptocurrency exchanges and see if I can find anything that
might be related to them.
I'll look at different sources where they might expose an identity that I can track.
I might be able to trace that to a wallet address.
Once I get a wallet address, I can go to the blockchain for that particular cryptocurrency,
and I can see what other transactions have been done.
If they're doing something that's related to a dark web form where they're selling
something in that wallet as a recipient, now I have another capillary path I can follow
out on the dark web.
And I can say, "Hey, what else is this person doing selling?
Who's buying?"
And I can follow those streams.
All of that information is going to be based on simply being told these guys are doing
some illicit transactions, go and track them down.
I need to use a browser.
I need to use a blockchain analysis tool.
I might need to install a particular wallet.
I might need social media tools like Telegram or Discord in order to track them across different
forms.
I need to access the dark web.
Our platform allows you to do all of that stuff, but those tools, when given to a 20-something
year-old analyst right out of school, can also be used to bypass policy guidelines inside
of the organization.
So the risk of abuse is huge.
So all of those things need to be audited and managed.
If the users abuse it and these organizations don't self-report or can't tell Congress
a good story, they don't get authorization to conduct those investigations anymore.
So Authenticate can enable someone to do all the things you just described.
So you're browsing the internet through and the internet's not actually touching your
machine, which sounds a bit like magic, although it works, right?
That's so hard.
And it allows you to use all those tools and save all those records.
That's right.
All at arm's length and all with a policy and audit framework to make sure people are
within compliance.
And you and your colleagues at Boozal, and without talking about the specific tools and
companies you work with, because I know you don't want to talk about that necessarily,
but solutions like these, what are the other kinds of solutions that you need to be an
effective O-center?
Yeah.
What are their technical solutions?
Absolutely.
So, to kind of tie the room together with a nice rug to use.
With a nice Lebowski reference.
Thank you.
With a nice Lebowski reference.
Absolutely.
Boozal and in executing this mission is using the full stack of whatever we need in order
to execute the mission successfully.
And in some cases, we procure that, right, commercially available data, for example, helps
us to identify the initial leads and triage and plan in terms of what this investigation
or what this open source collection is going to look like.
We will then use access tools or build access tools in many cases in order to achieve what
Scott had mentioned, which is getting into the environments that we need to get into.
And that requires the use of a variety of different access factors and capabilities.
Some of which Boozal and has built to kind of fill some of the niche gaps that we need
to fill.
And then from a processing and exploitation standpoint, Boozal and is one of the largest
providers of AI to the federal government, is now developing a lot of artificial intelligence
machine learning, automation, et cetera, to make truly quick and actionable intelligence
and insights out of the data that we're collecting.
And so, we're applying any tool that we need for the specific mission, kind of like a locksmith
with a bag of tools, understanding what that lock looks like, how it's designed, who built
it, where it was procured from, is a first step.
And so, we'll use whatever type of data and whatever type of automation to understand
that environment that we need to, we'll apply that from an access perspective.
Again, we'll use whatever that environment looks like, whatever makes us look most normal
and secures both our presence and our ex-filtration, and then, of course, on the processing.
So how many different, I know Boozal and does this a lot, you sort of buy what's out there
and then you build what isn't or what needs to be customized to whatever it is you're
doing.
Boozal and does this across the company, but sticking to open source intelligence technology.
Let's say one of your analysts is out there looking at Russia, China, whatever it is.
How many different providers of the software is he using throughout the course of, like,
we're talking two different pieces, two different companies, solutions he's working with, whether
it's companies like theirs or others, is it 10, is it 20, is it like, how complicated
is this and integrating different company solutions?
My answer is a terrible answer because it depends.
Well, then work on a better one before you give it to me.
Yeah, absolutely.
Well, I think, ultimately, that's a really important point, like, ultimately, understanding
the customer's mission and what the end goal of that mission is helps define what tools
or data commercially available or proprietary you may need, the tools, like, authenticates
that you may need to leverage to do that mission.
And so that really shapes, I think, what the overall solution needs to look like from a
data collection or prioritization of collection perspective.
And the tools or data sets you may need to accomplish that, because, ultimately, what
I find is a lot of the public sector customers need some very specific tailored output.
And they're coming to the market to get what they feel is the best fit, but it may not
ultimately be the truest fit to what they're really asking for.
And so, ultimately, the tool set and the data that go into those tools and the workflows
and the processes that create the tailored output really define what the overall solution,
to your point, needs to look like to really deliver that outcome.
That dovetails into what Scott had mentioned around the policy and the guidelines there.
There are a lot of cases where we are obligated to follow the guidelines in which the agency
or the organization that we're supporting has applied to us.
And so, in terms of authorities, in terms of auditing, in terms of providence of information,
in terms of compliance, like ICD compliance, for example, and the intelligence community
is huge.
And so that will often dictate the tools, the resources, the data, the technologies that
we implement.
But then what we're finding, what we have found in the last couple of years, is that open
source, to your earlier point, has truly evolved.
It's not your father's oscent.
And so, we're finding that there are a lot of new white space that we can apply where
policy doesn't necessarily exist yet, or the information environment has evolved so much
further and so much faster than where we used to be, where it's allowing companies like
Booz Allen to start to innovate in that space and influence, like, hey, this is what Wright
looks like, this is what we should be considering, and this is how it's actually going to impact
your mission moving forward.
May I add an example to that?
Because it's something that hasn't happened with another int previously, which is open
source information, as it becomes more mature and becomes sort of an int of a first resort,
as it is often called, is based on collections that are not necessarily classified or secure.
That allows a collaborative nature of the data that lets organizations exchange information
without violating any of the conditions by which they access the data.
It allows government organizations to share across partner governments without revealing
sources and methods.
There's this wonderful sort of collaborative potential of open source that lives below
the waterline because it is collected out from publicly available or commercially available
information that's going to foster this sort of triangulation that happens across multiple
endpoints.
You're absolutely right.
I think some of our proudest moments are when we collect a piece of information that a stakeholder
will look at and go, "This is sensitive.
This is classified."
And we go, "No."
It's not.
We found it in the trash bin at the internet.
Absolutely.
And it happens more than you think.
100%.
100%.
And the other, to your point, about collaboration is it makes information shareable, releasable,
cross-agency partners, where there's a gap or a historical gap between, let's say, the
intelligence community or the federal law enforcement community, but even with foreign
partners where information sharing is absolutely critical to ongoing relationships or operations
or things like that, when it's open source, it becomes shareable even when the notion
of this new information is it should be classified.
It's very sensitive.
So when you're working in this space, you yourselves, not just in the course of doing
open source investigations, but just by being solutions in this space become a target.
I imagine that all your companies have to deal with cyber penetration, cyber attempts,
some by criminal actors, some by state actors, without getting into anything too sensitive
on the details.
How do each of your organizations handle that?
I know for booze, it's a larger question, but especially for our founders here.
So we use a lot of our own intelligence and collection to look at ourselves, just like
we would apply it to a customer.
So we're sort of eating our own meal, if you will, to look at and leverage the information
that we're gathering focused on XeroFox itself.
Do you have an internal intelligence organization that focuses on counter threat?
Yes.
There's a sock and an internal team that leverages our own platform as well as other tools to
look at our own environment and be aware of what's going on out there, just like we would
apply the platform and the information to be applied to any other customer as well.
That way, we can also be aware of our own attack service and what threat vectors are
out there looking to probe, attack, scan, social engineer, plan attacks against the company,
its resources, or its supply chain.
Similarly to what Ryan described, we have a fully staffed set of resources inside of
the organization for security operations and organizational security, et cetera.
We follow a number of certification guidelines where FedRAMPs are certified.
We have ISO certifications, et cetera, et cetera.
There's also this outside-the-box thinking beyond just following the letter of law of
those certifications.
We have a pretty sophisticated security architecture team that reviews everything we do for vulnerability
at scale.
We've designed our systems to segregate information across multiple resources so that if we are
storing information on behalf of somebody, it's not stored in a single place and the
keys are split across various places.
You really want to look at your security posture with the expectation that something
will break.
How do you minimize the impact of that one thing breaking?
We go through that from a technical perspective, from a staffing perspective, from an operational
perspective.
I might be wrong on this, but I would guess the biggest target for your company is it's
the actual machines that are touching the internet on behalf of the client somewhere.
That's where virtualization comes in.
Everything we do is built on a container that's built for task in region, and when it's done,
it gets torn down, burned down.
We are able to re-inject state data into the next session that it's brought up.
It looks relevant for the purpose at hand, but that compute instance no longer exists.
That's a real nice way to keep those threats at bay.
The other thing to think about is what happens from the virtualization perspective, the execution
of the code in that instance lives on that machine, but the user is only seeing a streamed
remote display of that session.
There's no network traffic.
There's no data transiting modulo policies defined by an administrator.
There's no anything transiting to the end user's endpoint other than display data, which is
secure and benign.
What is open source technology that you think doesn't exist that needs to exist in the next
five years, especially as AI moves more to the center of everything?
I think there are certain technologies that are starting to exist.
I think we're in our relative nasancy with countering adversarial AI, from a misinformation
to disinformation perspective.
There are a lot of technologies that are starting to emerge out of that landscape, but I think
that that is one that's going to become more of a focus, and I think it's going to be more
prolific over the next couple of years, because the ability to create new information using
AI, realistic and believable AI is incredible.
There are certain information environments where information coming out of that environment
is intentional misinformation and disinformation, and so distinguishing between what is real
and what is not, I think, is going to be incredibly important.
Absolutely.
I think that's the landscape that's really going to emerge.
I'll tag team onto that, because I am more concerned about the ability of AI to lower
the threshold for actors to ramp up the velocity in the variety of tax, and I'm not as worried
about the technical to technical IT to IT attacks, because there's so many bright and
creative and dedicated people in the US working on those types of problems.
I'm more worried about, as you said, the ability for those threat actors to propagate attacks
that are more social engineering-based, to trick users or citizens of the world into
believing the misinformation or the various scams they can perpetrate, or the way they
can use social engineering to get access to a company or an agency, and then leverage
that and gain further access to do more harm.
If history is prologue, our track record as a society being skeptical is not particularly
good.
That's true.
Absolutely.
It's been so interesting to look at things like the emergence of some of these large
language models.
Let's say DeepSeq, for example, caught the world by surprise when it came out.
This is the Chinese large language model.
Exactly.
Exactly.
And watching how DeepSeq reacts to a prompt using its Western-facing version versus its
domestic version is pretty incredible.
And the reason is because the data that it's trained on is completely different.
And so I think there's some opportunity to look at things like that.
When you look at some of these models, some of these technologies, the data that it's
being trained on, you can start to see what the reaction is.
You can start to see what the different perturbations would be, the different permutations of an
output.
And then you can start to at least distinguish between what truth is versus what has been
curated in some of these particular cases.
Scott, you started authenticating, you talked about this earlier, originally to be a commercial
solution.
What was your vision for what that would be and why do you think that it didn't work
out that way, and you ended up focusing much more on governance?
One of the rare times I get to provide revisionist history, I guess.
I used to brief CISOs when we were doing the attempt to sell our product to a commercial
security organization.
And I would say, imagine 20 years ago, I could give you a magic box that would sit on every
employee's desk, and it would give them real-time information on all their competitors.
It would give them access to information.
They could answer any question about something that they were working on.
I could even wire applications into it, so if they wanted to do expenses or if they wanted
to book travel or they needed to buy office supplies or whatever, they could do it from
this magic thing on your desk, and you, Mr. Ciso, you didn't have to install a lick of
software in your environment.
Everybody would --
No.
No.
The only trade you have to make is you open your ports to the Internet and you allow arbitrary
code to transit the Internet and execute on everybody's endpoint.
Nobody in the room would take that trade, but that's exactly what the browser does.
That's exactly what you do when you use some of these other Internet-connected applications.
You're allowing arbitrary third-party code to transit your network and execute on your
endpoint.
The world of security is built around determining more quickly what happened and how do I remediate
as opposed to preventing, right?
It's the old Ben Franklin axon.
While that's true, it has become standard operating procedure.
It is the behavior.
It is the environmental variable that organizations plan around.
Internet-connected devices are part of the equation.
The idea we had for an enterprise-grade secure browser with managed access to third-party
websites and full audit and administrator control, I don't know that it would have worked
because the momentum of the alternative was so strong.
People don't want to change behavior patterns or change the user's expectations of what
they can get access to.
Sidebar, which I find very ironic, because things like VPNs for remote access are not
opt-in.
So we were thinking that we would build that browser that would solve that problem for
the organizations.
It turns out that the organizations didn't prioritize the way that we saw fit.
The irony of all this is that our first customers that pulled us into this managed attribution
model were government customers, which is in my previous company, when we sold to Google,
we had something like 40,000 businesses worldwide using us as a SaaS email processor.
That's why Google was interested in us.
We did not have a single government organization.
And yet a short, whatever, six years later, the government was an early adopter for some
of these SaaS capabilities for providing managed attribution resources for their people.
So that's pretty interesting.
What we've all been describing, or what you've all been describing, this is a real profession,
and it's one that takes a certain amount of technical mastery, tradecraft, understanding
skepticism and information.
It's pretty hard to find that in one person.
In a sense, what I'm hearing from you is it actually might be harder to find the ideal
ocentor than it is to find the ideal intelligence analyst working on only classified systems,
because there's all these other considerations involved.
Does this create, from your different perspectives, what's the training pipeline issue?
What are the bottlenecks?
Yeah, that's a fantastic question.
It's something that we've wrestled with at Booz Allen for a couple of years, and really
trying to transform the way that we are looking at and applying ocent, at least from our national
security perspective, where we have consolidated, unified, if you will, both our capabilities,
but also our tradecraft and our methodologies.
And what that leads to is the talent pipeline required to execute those missions.
And so one of the things that we've really placed our focus on is the technical mastery.
And so I think that we can take and make or create a successful ocent collector, but again,
we're moving back from the ocent analyst vernacular and moving into ocent collector, ocent exploitation.
And so we can start with somebody that understands the information environment, has an IT background,
understands networking, understands network protocols, things like that, and teach them
certain tradecraft as it relates to human-like collection, if you will.
On the flip side, we can take somebody that is a traditional collector from an operational
standpoint and upscale them with some of the technology mastery that they need to understand
what they look like and how to operate in these environments.
And I think the third piece that's really critical, as I think Scott alluded to earlier, is cultural
and a language understanding.
There are a lot of really sophisticated translation tools that are out there, but in some of these
places understanding the culture and the environment of the landscape that you're traversing and,
of course, being able to read, understand, synthesize the language as you're operating
I think is pretty critical, too.
And technology can help with that, but it's not the end solution.
And so we can take one of those three legs of the bar stool and augment in the other
two where possible, but you're right.
It is like finding a unicorn these days because those three things are starting to converge
between traditional oscent collection, cyber, and then, of course, AI when it gets applied
on top of that.
Yeah.
I think a lot of it comes down to the tradecraft and the process and the culture you create
around the diligence to go work at that problem and understand where the customer needs to
get to and how you need to fit into that model to get them there.
And understanding and building the tradecraft, and which could be, as you said, any number
of learning courses, tool training, documentation, they may need to review to get up to speed.
We've had a lot of success finding really enthusiastic kids out of college, former law
enforcement, former DOD, but there is always an aspect of getting them onto the tradecraft.
And to our earlier discussion, some customers want things you deliver out of the box, which
is great, because that's what you've designed the solutions around.
But there are customers that need tailored outputs, tailored deliverables, tailored reports
based on their PIRs and RFIs and things like that.
And so you really have to hone in on the general concept of what you will be doing in open
source or deep dark web collection, how that translates into a deliverable, and then the
tradecraft around that from a tool or training perspective to get them there, to be able
to go let them loose, to go help protect that customer or the country.
If you look at the proposed Authorization Acts, there is a desire to codify practices
and training and standardize on tools and standardize on collaboration channels and
things like that.
Those things take a while to develop.
It's fine to put the directives out there.
It's fine to put the money in place, but the organizations, it's like turning a big ship.
The organizations are going to be slower.
OSINT needs to be a job track.
We started to see that in some of the large organizations that we work with.
If you're selling into an enterprise organization, the SOC/NOC analyst was kind of an OSINT-er
to begin with, because it's the crafty network person that was out there trying to find vulnerabilities
in systems or things like that.
If you move into the other disciplines outside of CTI, into pure OSINT or, say, financial
fraud investigations, things like that, those same genetic predispositions need to be in
place.
People need to be inquisitive, curious, and all of that.
There's a huge impetus for us to make sure our people, our users, are trained and capable,
because we're giving them a set of tools that allow them to hurt themselves and hurt us.
If somebody orders a pizza to their base, which has happened using the managed attribution
platform, we have to burn down a whole bunch of infrastructure, rotate IPs, and do all these
other things, because that's now a signal out there.
There's not only a requirement that the customer gets smarter, but we as a vendor also need
to invest in making sure our users are smarter so that they don't hurt themselves or us.
Absolutely.
It's an educational journey, not just for our own people, but also for the people that
we're supporting to show them exactly what Wright looks like from a tradecraft from an
expertise perspective, and that is a cultural shift that we've seen over the last couple
of years.
That's a great point.
The ability to, using tools or other data sets to go create and gather information also
requires you to train the end user or the customer on how it could be used, how it should
be used, and maybe what you shouldn't be doing with it.
Definitely what you shouldn't do with pizza.
Not ordering pizza.
Thank you for listening to this episode of Cogs of War.
Stay safe and stay healthy.
[end of transcript]
Podcast Summary
Key Points:
Open source intelligence has evolved to rely on specialized tools and platforms for data collection and analysis.
The podcast episode features discussions with leaders from private sector companies dealing with open source intelligence.
Companies like XeroFox offer software platforms for gathering open source data, including deep and dark web information.
Summary:
The Cogs of War podcast episode delves into the evolution of open source intelligence, highlighting the shift towards using specialized tools and platforms for data collection and analysis. The episode features insights from industry leaders like Ryan Curran from XeroFox, who discuss the importance of leveraging software platforms to gather various types of open source data. These platforms enable organizations to gain a holistic perspective on potential threats and prioritize actions based on AI-driven analyses.
The discussion emphasizes the challenges of navigating complex information environments, the need for human verification in data analysis, and the technical tradecraft required for successful open source intelligence operations. Additionally, real-world examples illustrate how companies like XeroFox assist public sector organizations in monitoring and addressing threats such as scams and fraudulent activities targeting executives and high-profile individuals.
FAQs
Open source intelligence has evolved from simple web scraping to using specialized tools and services to analyze and protect data.
The guest started in Silicon Valley in the late '80s, worked at Apple Computer, and eventually moved into product management and founding companies.
The guest was presented with an opportunity to apply their technology to managed attribution, which led them into cybersecurity and open source intelligence.
Booz Allen leverages technology to focus on open source collection and exploitation, shifting towards technology enablement and technical considerations.
Open source intelligence has become more complex due to diverse online environments, increased data on the deep web, and higher security measures requiring technical expertise.
Automation is crucial for gathering data at scale, but good collection and technical tradecraft still require human expertise to navigate the complex information environment.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.