Overview
The Databricks integration lets Cotool agents query security data in your Lakehouse through a Databricks SQL warehouse. It provides:- Data discovery — browse Unity Catalog catalogs, schemas, and tables, search for tables by name, and inspect a table’s columns, partition and clustering columns, comments, and sample rows.
- Read-only SQL — run Databricks SQL for investigations, alert triage, threat hunting, and code detections. Results include typed columns and a
truncatedflag when more rows matched than were returned.
Before you start
You will need:- A Databricks workspace with Unity Catalog enabled. Legacy
hive_metastoretables also work; see Legacy Hive metastore tables. - Workspace admin rights, to create a service principal and generate an OAuth secret for it.
- A SQL warehouse Cotool can use. We recommend a dedicated serverless SQL warehouse:
- Serverless warehouses start in seconds. Classic and pro warehouses can take several minutes to start after auto-stop, and Cotool’s first queries after an idle period may time out while they start.
- A dedicated warehouse keeps Cotool’s queries and their cost separate from your other workloads. Every Cotool query runs on your warehouse and consumes DBUs. Set an auto-stop interval that suits your budget.
- The owner of the catalogs or schemas holding your security data (or a metastore admin), to grant read access.
Cotool queries through a SQL warehouse, not an all-purpose cluster. If the
HTTP path you have starts with
sql/protocolv1/o/, it belongs to a cluster;
create or pick a SQL warehouse instead.Set up Databricks
1
Create a service principal
- In your workspace, click your username in the top bar and open Settings > Identity and access.
- Next to Service principals, click Manage, then Add service principal > Add new.
- Enter a recognizable name (for example
cotool) and click Add. - Open the service principal and copy its Application ID (a UUID). This is the Client ID in Cotool.
2
Give it Databricks SQL access
On the service principal’s Configurations tab, make sure Databricks SQL access is checked, and click Update.
This entitlement is easy to miss. Without it, the service principal cannot use any SQL warehouse, even with Can use permission on one, and Cotool’s connection check fails with a permission error.
3
Generate an OAuth secret
- On the service principal, open the Secrets tab and click Generate secret.
- Choose a lifetime (up to 730 days) and click Generate.
- Copy the Secret value immediately. Databricks shows it only once.
4
Grant access to a SQL warehouse
- Open SQL Warehouses and select the warehouse Cotool should use.
- Click Permissions, add the service principal, choose Can use, and click Add.
- On the Connection details tab, copy the Server hostname. Copy the HTTP path (for example
/sql/1.0/warehouses/1234567890abcdef) only if you want to pin this warehouse.
5
Grant read access to your security data
In the SQL editor (or Catalog Explorer > Permissions), grant the service principal read access to the catalog that holds your security data. In Privileges granted on a catalog apply to every schema and table in it, including ones created later. To limit Cotool to specific data, grant
GRANT statements, refer to a service principal by its Application ID:USE SCHEMA and SELECT on individual schemas or tables instead. USE CATALOG on the parent catalog is always required.Do not grant MODIFY, ALL PRIVILEGES, or ownership. Cotool only needs to read.Connect in Cotool
- In Cotool, go to Platform > Integrations > Databricks and click Connect.
- Enter the values you collected:
- Click Connect. Cotool exchanges the credentials for a token and reads (or picks) the SQL warehouse to confirm both work before saving. If the warehouse is not serverless, the connect form shows a warning, because its first query after auto-stop waits for it to start. If the check fails, the error explains which part to fix (see Troubleshooting).
Cotool only sends credentials to Databricks-operated domains, including the GovCloud and Azure Government and China variants. Custom vanity domains are not supported; use the workspace’s Databricks hostname.
Personal access token (fallback)
If your organization cannot create service principals, choose Personal access token and paste a token instead of the client ID and secret. Generate it for a dedicated, least-privilege identity rather than a person’s account, because Cotool can read whatever that identity can. Tokens expire, and workspace admins can disable them; when that happens Cotool’s queries start failing with HTTP 401 until you paste a new token.Network access
If your workspace restricts inbound access, Cotool must be able to reach it:- IP access lists — allow Cotool’s egress IP addresses, listed on the Databricks connect form in Cotool (Platform > Integrations > Databricks > Connect).
- Private Link with public access disabled — Cotool cannot reach the workspace. Keep a public front-end endpoint restricted by IP access list instead.
HTTP 403 (IP access list) or a connection failure before any response (Private Link or firewall).
How Cotool queries Databricks
- Discovery is incremental. Agents list catalogs, then schemas, then tables, and describe only the tables they need. Discovery uses the Unity Catalog metadata APIs, which do not wake the SQL warehouse. Sampling rows does use the warehouse.
- Results are capped. Queries return 500 rows by default and at most 10,000, with a 4 MiB result cap. When more rows match, the response says
truncated: true, and the agent narrows or aggregates the query. - Queries time out and are cancelled. Each statement has a timeout (300 seconds by default, at most 30 minutes). Cotool cancels statements that exceed it, so abandoned queries do not keep running on your warehouse.
- Queries are attributable. Cotool’s statements appear in Databricks Query History under the service principal, so you can audit and cost them.
- Partition or liquid-cluster large log tables on an event date or timestamp column. Cotool shows agents these columns and tells them to filter on them first.
- Add table and column comments describing each source (for example, “AWS CloudTrail management events”). Agents see comments during discovery and use them to pick the right table.
- Use one table per log source rather than one mixed table, so agents do not have to filter by source type.
Environment map
After you connect, and weekly after that, Cotool maps the workspace’s Unity Catalog tables: each table’s columns, partition and clustering columns, a few sample rows, and a summary of what it holds (for example, which log source it is). Agents start from this map instead of rediscovering the catalog, and detection planning uses it to know which log sources you have. View it on the Databricks integration page in Cotool.- The map covers the tables the service principal can see, up to 300. It skips the
system,samples, andhive_metastorecatalogs andinformation_schemaschemas; agents can still query those directly. - Listing tables uses the Unity Catalog APIs and does not wake the warehouse. Sampling runs one
SELECT * ... LIMIT 3per table, so mapping briefly starts the warehouse. If the warehouse cannot start, Cotool maps tables without sample rows. - Tables created between runs are not in the map until the next run, but agents can still find them with discovery.
Legacy Hive metastore tables
Tables in the workspace-localhive_metastore catalog are not served by the Unity Catalog APIs, so Cotool discovers them with SHOW and DESCRIBE statements on the SQL warehouse. This works, with these differences:
- Table-name search across the catalog is not available; agents browse schema by schema.
- Discovery wakes the SQL warehouse.
- Access is controlled by legacy table access control rather than Unity Catalog grants. Grant the service principal
USAGEandSELECTon the relevant databases, or upgrade the tables to Unity Catalog.
Troubleshooting
Local development
Populatecogent-backend/.env with DATABRICKS_HOST, DATABRICKS_HTTP_PATH, DATABRICKS_CLIENT_ID, and DATABRICKS_CLIENT_SECRET, then run npm run bootstrap-tools <API_KEY> from the repo root.