Some of the more useful conversations about AI and data quality lately haven’t come from AI-native companies. They’ve come from data companies old enough to have lived through several previous waves of technological change and simply kept collecting.
That was the setup for a recent episode of The Innovation Blueprint Podcast, ORIL’s series on emerging PropTech, hosted by ORIL’s CEO Roman Havrylyuk. His guest, David Lovins, is CEO of The Warren Group, which has been collecting public records data since 1872 and has been led by David for the past three decades. The conversation moved through data standardization, AI, and startup advice, but it’s less a profile of one company than a look at a set of problems that keep showing up across property data generally — which is really what these episodes are for.
Roman opened with his own reaction to the company’s age, and it set up the real question behind the episode: “just when I first learned about the age of the company, I was really surprised — being at that phase and maintaining the technology realities is quite impressive.” Surviving a hundred and fifty years of change is one thing. Staying technically current through the current one is a different, harder question, and it’s the one the rest of the conversation actually tries to answer.
What “Understanding the Data” Means in Practice
The Warren Group runs 15 to 20 different real estate datasets — deed, mortgage, bankruptcy, building permits, rental data — and serves roughly 40 to 50 industries connected to real estate, from proptech and fintech to insurtech and home services.
David’s framing of what separates a data company from a data reseller is worth pulling out on its own: “we really understand where the data’s coming from, how to collect the information, how to synthesize it, how to normalize it.” His broader point is that a lot of the market sells records without a working understanding of how that data was actually produced, and that gap tends to surface later, when a customer asks a specific question the vendor can’t actually answer. It’s a useful test for evaluating any data provider, not just this one: did they build the pipeline, or are they passing along someone else’s feed?
Roman connected that to something he’s watched shift across the wider market: “the overall market awareness also evolves over time. The straightforward usage of any data has become a commodity, I guess, and now more and more companies are learning that there’s this data out there — and with the use of AI, there’s much more they can do with it.” That’s less a Warren Group observation than a description of where the whole industry is right now: raw access stopped being the hard part a while ago.
What AI Actually Changed, Versus What It Just Made Visible
Ask David how AI has changed the business and two different answers come out — one about customers, one about internal operations.
On the customer side: a wave of startups, now roughly 30 to 40% of new prospects by his estimate, are showing up wanting to feed Warren Group data straight into an AI engine to build a new application or analytics layer. Most are bootstrapped, and most don’t ask for a single dataset — they ask for seven, eight, nine at once, then have to figure out how to combine them. David has been building some of these himself: “I’ve done a lot of vibe coding and learned how very easily it is to put an application together — a real true working prototype that can be promoted out there as a custom solution.”
Roman flagged this as a trend he expects to keep accelerating, not level off: “I’m already seeing, and foreseeing, a lot more of that coming — where a lot more early-stage startups will have more access to the tools they need to build out early-stage products.” Internally at Warren Group, AI shows up less as a headline feature and more as infrastructure spread across the business — standardizing incoming data, surfacing insights, and doing quality control across sales, customer service, and collection processes.
How Data Vendors Are Handling the Early-stage Market Differently
One thing that came up at length: how a data vendor decides who’s worth working with early. Bootstrapped startups often ask for more data than they can reasonably afford or use, and David described a practice of scoping that down rather than selling to the ask — starting with a smaller geography as a proof of concept, for instance, instead of a nationwide dataset from day one.
“We try to work with them not to over-leverage them,” he said. “We try to work with them like a real customer, part of the family, part of the company, to make sure they’re successful.” He drew a direct contrast with a more transactional approach he sees elsewhere in the market — quote a price for whatever’s requested, move on.
Roman’s take from ORIL’s side was similar in spirit but stated as a general principle rather than praise for one vendor: “I believe that building partnerships and helping grow the network around your business is the way to go. Working with larger businesses that can afford the actual cost of the service is generally the preferred way, but if you can grow with a company, you’re growing the next generation of enterprises that you can then grow together with.”
Awareness Is Catching Up to What’s Actually Available
A theme that ran through the second half of the conversation: customers increasingly don’t know what data exists until someone shows them. David called it “keeping up with the Joneses” — most prospects want more data, they just don’t know it’s out there until they’re in a conversation about it. When Warren Group doesn’t have the right dataset itself, David said the company will go find it through a partner rather than leave the customer to sort that out on their own, which is part of why he positions the company as a solutions provider rather than a pure data seller.
This is where Roman brought in something ORIL sees directly in its own client work, not just as an observation about the guest’s business: “enriching the existing data sets from maybe various data providers — that’s something that we hear often as well lately.” It’s a specific, practical problem: a client has one data source that’s close to what they need but not complete, and the fix isn’t a bigger dataset, it’s connecting several partial ones into something coherent.
The Problem Every Data Conversation on This Podcast Keeps Circling Back To
If there’s a theme that keeps surfacing across conversations with property data companies, it’s this one: none of the underlying datasets were built to talk to each other, and stitching them together after the fact is genuinely hard.
David named it directly as a near-term priority: “since we have a number of different data sets, some of them don’t talk to one another… creating a unique identifier across all data sets and making it easier for the customer to ingest the data is going to be critical.” His reasoning for why this matters more now, not less, is worth sitting with — hand an AI system 15 disparate databases and ask it to reconcile them, and it may produce something that looks right without actually understanding the nuance of how each dataset was structured or where its blind spots are.
That connects to a broader point about data quality David raised: public records data is locally sourced, inconsistently formatted, and never fully standardized out of the gate. AI helps here too, but in a specific way — using it internally to catch anomalies in the data before a customer does, rather than assuming AI will clean things up on its own. And when the data isn’t perfect (it never fully is), David argued the more important move is transparency: telling the customer exactly where the known gaps and risks sit, so they can adjust their own models accordingly, instead of finding out the hard way.
What to Keep in Mind for Startups Before They Build Anything
Asked for advice for early-stage companies building in real estate tech, David didn’t start with technology. He started with a question: do you actually understand the end game — the mission of what you’re building — before you start pulling data?
His broader point pushes back on a narrative he’s hearing more of: that AI alone can collect and interpret real estate data without real domain understanding. “That is not the case,” he said. “It is very hard to understand real estate. It’s not rocket science, but you really gotta understand what data you really want, and understand it, before you build a product.“ He offered a plain metaphor for where his company fits in that process: “we’re the gas — you can build your Ferrari, your Toyota Camry, or a plane, but ultimately we’re the fuel that runs it. You have to understand what type of fuel you need before you build the car.”
Roman closed that thread with an observation of his own about how the problem itself has shifted, not just the tools: “before, let’s say, two years ago, what we’ve noticed is the lack of knowledge — the exposure to what’s available. Now it’s the opposite, where people are just bombarded with the capabilities, especially with AI, and they need to be able to choose which route they want to go — narrowing the focus into something specific and understanding what part of the service you want to provide to your customers.” It’s a fair summary of the whole episode: the constraint used to be access, now it’s judgment.
Where ORIL Fits In This Picture
None of the problems in this conversation are unique to one data vendor. Datasets that don’t talk to each other, records with inconsistent formatting and no reliable identifier, an AI layer that needs clean and well-understood inputs to be trustworthy — these are exactly the kind of problems ORIL works through with real estate and proptech companies: data integration across disparate sources, data enrichment that fills the gaps a single dataset can’t cover on its own, and visualization that turns a normalized dataset into something a team can actually act on, not just query.
If you’re a startup weighing which data provider to build on, or an established real estate company trying to make several data sources work together reliably, that’s the layer worth bringing in a technical partner for early — not after the data problems have already shaped the product.
Join the Conversation
Want to join the podcast? Drop us a line and share your insights on how innovation is shaping the future of real estate and PropTech!