CompStak AI is here! Pair powerful AI capabilities with CompStak’s trustworthy data. Click to learn more! CompStak AI is here! Click to learn more!
Help us direct you to the right place to sign up

A commercial real estate ontology is the model behind property data. It defines what kinds of entities exist (property, tenant, lease, landlord, transaction) and how they relate to one another. It is not the data itself, and it is not the process of matching duplicate records. It is the layer of meaning that lets a pile of documents behave like a connected system.

Key Takeaway: An ontology and entity resolution are two different layers, and conflating them is the most common mistake in this space. The ontology defines the types of things and how they relate. A lease connects a tenant to a property for a term. A sale connects a buyer and a seller. A landlord owns properties. Entity resolution is a separate job. It decides which records point to the same real-world thing, such as whether “123 Main St. LLC” and “Main Street Holdings” are one landlord. You need both. But the ontology is what makes data queryable. Once documents are tagged to it, scattered PDFs become a structured system that analytics can run on directly.

What a Commercial Real Estate Ontology Actually Is

An ontology assigns structure to meaning. A “tenant” is a distinct kind of entity. A “lease” is a relationship between a tenant and a property over time. A “landlord” owns one or more properties. Those definitions stay stable even when the raw data is messy. The point is not to store the lease. The point is to know what a lease is and what it connects, so any lease from any source drops into the same model.

The term comes from computer science and information science, where an ontology formally defines the entities in a domain and the relationships among them. That makes it different from a taxonomy, which is only a hierarchical classification. The W3C’s work on ontologies in the semantic web is the closest analog outside real estate. The goal there is the same one it is here: make disparate data machine-interpretable by defining what each thing is and how it connects to everything else.

Commercial real estate is a hard domain to model this way. Addresses change. Landlord entities get restructured for tax and financing reasons. Brokers abbreviate tenant names differently on every abstract. A useful CRE ontology holds a consistent picture of properties, tenants, landlords, and the deals that link them. It does that no matter how each source labeled the pieces.

Ontology, Entity Resolution, and Schema Are Three Different Things

These three terms get used interchangeably, and they should not be.

A database schema defines the technical structure that holds data: tables, fields, and data types. A schema might have a “tenant_name” field. It says nothing about whether the values in that field mean anything consistent.

An ontology defines meaning and relationships, independent of storage. It says what a tenant is and how a tenant relates to a lease and a property.

Entity resolution is the process that populates the ontology with clean entities. It decides that “Acme Corp,” “Acme Corporation,” and a regional subsidiary filing are one tenant, or that two LLCs trace back to one institutional landlord.

The distinction is practical, not academic. Entity resolution without an ontology gives you cleaner records with no structure to query. An ontology without entity resolution gives you a well-designed model fed by duplicates. And two datasets can share an identical schema yet be unusable together if their resolved entities disagree, because joining them multiplies noise instead of insight.

The Real Value: Turning Unstructured Documents Into a Queryable System

Here is the part that gets lost. The payoff of an ontology is not a marginally better number at the end. It is a transformation at the beginning.

Most commercial real estate data starts as documents. A lease is a PDF. A sale is a filing. An abstract is a broker’s notes. On their own, these are unstructured. You cannot ask them a question. Tag each document’s facts to an ontology, though, and something changes. “This lease relates this tenant to this property for this term at this rent” stops being text and becomes a record in a connected model. Do that across a whole corpus and the pile becomes a structured dataset you can query as a whole.

That is the shift that separates a document archive from an analytics system. The companies that pull ahead on data do not simply collect more of it. They model it so it can be queried as a connected whole, and then everything downstream runs on that structure. The ontology is what makes the raw material usable. It is the difference between a folder of leases and a system that can tell you how much space a tenant occupies across a market.

How CompStak Turns Comp Documents Into Structured Data

CompStak’s raw input is exactly that unstructured pile. Lease and sale comps come from a network of 40,000+ verified CRE professionals through CompStak Exchange. Most arrive as abstracts and documents, not clean rows.

Turning that into a usable dataset takes two steps. First, verification. Every comp passes through a multi-step process before it lands in the platform. Machine learning models flag statistical anomalies, then a team of CRE data analysts reviews the results. Second, resolution. CompStak matches each verified comp against a resolved set of property, tenant, and landlord records instead of adding a standalone row. A tenant leasing under a slightly different name, or a landlord restructured since its last deal, reconciles to the right underlying entity rather than spawning a duplicate.

On top of that resolved foundation, CompStak builds relationships. Because tenant and property entities are resolved, the platform can link a tenant’s leases across time and buildings, which is what powers tenant renewal and relocation tracking. That is a concrete, shipped example of ontology work: resolve the entities, then relate them.

What This Means If You’re Aggregating CRE Data Yourself

For any investor, lender, or asset manager pulling data across multiple sources or a large portfolio, the information starts as documents, and documents do not answer questions. The problem is never a shortage of data. It is the absence of a structure that turns the data into something you can query.

That is the practical takeaway. The value is in the transformation, not in owning yet another spreadsheet. An acquisition underwriting, a loan-collateral review, and a peer-portfolio benchmark all depend on one thing. They need entities resolved correctly and related in a consistent model. Get that right once, upstream, and every downstream metric inherits it.

You can build that discipline in-house, or you can start from a dataset where the work is already done. Either way, the lesson holds: structure first, analytics second.

Frequently Asked Questions

What is commercial real estate ontology?

It is the model that defines what kinds of things exist in property data (property, tenant, lease, landlord, transaction) and how they relate to one another. It is the structure of meaning, not the data itself and not the process of cleaning it.

How is an ontology different from entity resolution?

They are separate layers. An ontology defines the entities and relationships. Entity resolution decides which records refer to the same real-world entity. The ontology gives you a model to query; entity resolution fills that model with clean, non-duplicated records.

How is an ontology different from a database schema?

A schema defines the technical structure of tables and fields. An ontology defines the meaning of the data and the relationships between real-world entities, independent of how it is stored.

Why does an ontology actually matter?

Because it turns unstructured documents into a queryable system. Once each lease or sale is tagged to the ontology, a pile of PDFs becomes a connected dataset. Analytics can run on it directly, instead of leaving you a folder no one can question.

How does CompStak use an ontology?

CompStak verifies each contributed comp with machine learning and analyst review. It resolves the comp to consistent property, tenant, and landlord entities, then relates them, for example by linking a tenant’s renewals and relocations across leases and buildings.

Does resolving entities affect comp accuracy?

Yes. Unresolved identities can split one lease into two comps or merge two tenants into one, distorting rent, concession, and WALT figures. Resolution upstream is what keeps those metrics honest.

Is CRE ontology only relevant to software companies?

No. Any investor, lender, or asset manager combining data across sources or a large portfolio faces the same document-to-structure problem. It is a data-governance issue, not a software one.

Where does this show up in CompStak?

In property, tenant, and landlord-level search across the platform, and in relationship features such as tenant renewal and relocation links, all built on the resolved comp dataset.

Related Posts

The Top 10 Office Sectors in Commercial Real Estate to Watch in 2025

The Top 10 Office Sectors in Commercial Real Estate to Watch in 2025

Net Effective Rent vs. Asking Rent: Why the Gap Matters for CRE Underwriting

Net Effective Rent vs. Asking Rent: Why the Gap Matters for CRE Underwriting

Understanding the strength of the industrial sector

Understanding the strength of the industrial sector