Skip to main content

Overview

The Databricks integration lets Cotool agents query security data in your Lakehouse through a Databricks SQL warehouse. It provides:
  1. Data discovery — browse Unity Catalog catalogs, schemas, and tables, search for tables by name, and inspect a table’s columns, partition and clustering columns, comments, and sample rows.
  2. Read-only SQL — run Databricks SQL for investigations, alert triage, threat hunting, and code detections. Results include typed columns and a truncated flag when more rows matched than were returned.
Cotool connects through the Databricks REST APIs (SQL Statement Execution and Unity Catalog) as a service principal using OAuth machine-to-machine credentials. A personal access token is supported as a fallback. Cotool only needs to read from Databricks. Databricks SQL has no read-only session mode, so the service principal’s grants are what keep Cotool read-only: grant only what is listed below.

Before you start

You will need:
  • A Databricks workspace with Unity Catalog enabled. Legacy hive_metastore tables also work; see Legacy Hive metastore tables.
  • Workspace admin rights, to create a service principal and generate an OAuth secret for it.
  • A SQL warehouse Cotool can use. We recommend a dedicated serverless SQL warehouse:
    • Serverless warehouses start in seconds. Classic and pro warehouses can take several minutes to start after auto-stop, and Cotool’s first queries after an idle period may time out while they start.
    • A dedicated warehouse keeps Cotool’s queries and their cost separate from your other workloads. Every Cotool query runs on your warehouse and consumes DBUs. Set an auto-stop interval that suits your budget.
  • The owner of the catalogs or schemas holding your security data (or a metastore admin), to grant read access.
Cotool queries through a SQL warehouse, not an all-purpose cluster. If the HTTP path you have starts with sql/protocolv1/o/, it belongs to a cluster; create or pick a SQL warehouse instead.

Set up Databricks

1

Create a service principal

  1. In your workspace, click your username in the top bar and open Settings > Identity and access.
  2. Next to Service principals, click Manage, then Add service principal > Add new.
  3. Enter a recognizable name (for example cotool) and click Add.
  4. Open the service principal and copy its Application ID (a UUID). This is the Client ID in Cotool.
2

Give it Databricks SQL access

On the service principal’s Configurations tab, make sure Databricks SQL access is checked, and click Update.
This entitlement is easy to miss. Without it, the service principal cannot use any SQL warehouse, even with Can use permission on one, and Cotool’s connection check fails with a permission error.
3

Generate an OAuth secret

  1. On the service principal, open the Secrets tab and click Generate secret.
  2. Choose a lifetime (up to 730 days) and click Generate.
  3. Copy the Secret value immediately. Databricks shows it only once.
On Azure Databricks, the secret must be generated here, in Databricks. A Microsoft Entra ID app registration’s client secret does not work, even for an Entra ID-managed service principal.
Secrets expire. Before then, generate a new one and update it in Cotool under Platform > Integrations > Databricks > Edit.
4

Grant access to a SQL warehouse

  1. Open SQL Warehouses and select the warehouse Cotool should use.
  2. Click Permissions, add the service principal, choose Can use, and click Add.
  3. On the Connection details tab, copy the Server hostname. Copy the HTTP path (for example /sql/1.0/warehouses/1234567890abcdef) only if you want to pin this warehouse.
If you leave the HTTP path blank in Cotool, Cotool picks a warehouse the service principal can use: serverless first, then a running warehouse, then pro over classic. If that warehouse is later deleted, Cotool switches to another one. Pin a warehouse when you need Cotool’s queries and their cost on a specific one.
5

Grant read access to your security data

In the SQL editor (or Catalog Explorer > Permissions), grant the service principal read access to the catalog that holds your security data. In GRANT statements, refer to a service principal by its Application ID:
Privileges granted on a catalog apply to every schema and table in it, including ones created later. To limit Cotool to specific data, grant USE SCHEMA and SELECT on individual schemas or tables instead. USE CATALOG on the parent catalog is always required.Do not grant MODIFY, ALL PRIVILEGES, or ownership. Cotool only needs to read.

Connect in Cotool

  1. In Cotool, go to Platform > Integrations > Databricks and click Connect.
  2. Enter the values you collected:
  1. Click Connect. Cotool exchanges the credentials for a token and reads (or picks) the SQL warehouse to confirm both work before saving. If the warehouse is not serverless, the connect form shows a warning, because its first query after auto-stop waits for it to start. If the check fails, the error explains which part to fix (see Troubleshooting).
Workspace hosts look like this: Cotool only sends credentials to Databricks-operated domains, including the GovCloud and Azure Government and China variants. Custom vanity domains are not supported; use the workspace’s Databricks hostname.

Personal access token (fallback)

If your organization cannot create service principals, choose Personal access token and paste a token instead of the client ID and secret. Generate it for a dedicated, least-privilege identity rather than a person’s account, because Cotool can read whatever that identity can. Tokens expire, and workspace admins can disable them; when that happens Cotool’s queries start failing with HTTP 401 until you paste a new token.

Network access

If your workspace restricts inbound access, Cotool must be able to reach it:
  • IP access lists — allow Cotool’s egress IP addresses, listed on the Databricks connect form in Cotool (Platform > Integrations > Databricks > Connect).
  • Private Link with public access disabled — Cotool cannot reach the workspace. Keep a public front-end endpoint restricted by IP access list instead.
Blocked connections show as HTTP 403 (IP access list) or a connection failure before any response (Private Link or firewall).

How Cotool queries Databricks

  • Discovery is incremental. Agents list catalogs, then schemas, then tables, and describe only the tables they need. Discovery uses the Unity Catalog metadata APIs, which do not wake the SQL warehouse. Sampling rows does use the warehouse.
  • Results are capped. Queries return 500 rows by default and at most 10,000, with a 4 MiB result cap. When more rows match, the response says truncated: true, and the agent narrows or aggregates the query.
  • Queries time out and are cancelled. Each statement has a timeout (300 seconds by default, at most 30 minutes). Cotool cancels statements that exceed it, so abandoned queries do not keep running on your warehouse.
  • Queries are attributable. Cotool’s statements appear in Databricks Query History under the service principal, so you can audit and cost them.
Agents query fastest when security tables are laid out for time-bounded reads:
  • Partition or liquid-cluster large log tables on an event date or timestamp column. Cotool shows agents these columns and tells them to filter on them first.
  • Add table and column comments describing each source (for example, “AWS CloudTrail management events”). Agents see comments during discovery and use them to pick the right table.
  • Use one table per log source rather than one mixed table, so agents do not have to filter by source type.

Environment map

After you connect, and weekly after that, Cotool maps the workspace’s Unity Catalog tables: each table’s columns, partition and clustering columns, a few sample rows, and a summary of what it holds (for example, which log source it is). Agents start from this map instead of rediscovering the catalog, and detection planning uses it to know which log sources you have. View it on the Databricks integration page in Cotool.
  • The map covers the tables the service principal can see, up to 300. It skips the system, samples, and hive_metastore catalogs and information_schema schemas; agents can still query those directly.
  • Listing tables uses the Unity Catalog APIs and does not wake the warehouse. Sampling runs one SELECT * ... LIMIT 3 per table, so mapping briefly starts the warehouse. If the warehouse cannot start, Cotool maps tables without sample rows.
  • Tables created between runs are not in the map until the next run, but agents can still find them with discovery.

Legacy Hive metastore tables

Tables in the workspace-local hive_metastore catalog are not served by the Unity Catalog APIs, so Cotool discovers them with SHOW and DESCRIBE statements on the SQL warehouse. This works, with these differences:
  • Table-name search across the catalog is not available; agents browse schema by schema.
  • Discovery wakes the SQL warehouse.
  • Access is controlled by legacy table access control rather than Unity Catalog grants. Grant the service principal USAGE and SELECT on the relevant databases, or upgrade the tables to Unity Catalog.

Troubleshooting

Local development

Populate cogent-backend/.env with DATABRICKS_HOST, DATABRICKS_HTTP_PATH, DATABRICKS_CLIENT_ID, and DATABRICKS_CLIENT_SECRET, then run npm run bootstrap-tools <API_KEY> from the repo root.