> ## Documentation Index
> Fetch the complete documentation index at: https://docs.cotool.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Databricks Integration

> Connect Cotool to a Databricks SQL warehouse for read-only investigation, detection, and hunting queries over Lakehouse security data

## Overview

The Databricks integration lets Cotool agents query security data in your Lakehouse through a **Databricks SQL warehouse**. It provides:

1. **Data discovery** — browse Unity Catalog catalogs, schemas, and tables, search for tables by name, and inspect a table's columns, partition and clustering columns, comments, and sample rows.
2. **Read-only SQL** — run Databricks SQL for investigations, alert triage, threat hunting, and code detections. Results include typed columns and a `truncated` flag when more rows matched than were returned.

Cotool connects through the Databricks REST APIs (SQL Statement Execution and Unity Catalog) as a **service principal** using OAuth machine-to-machine credentials. A personal access token is supported as a fallback.

Cotool only needs to read from Databricks. Databricks SQL has no read-only session mode, so the service principal's grants are what keep Cotool read-only: grant only what is listed below.

## Before you start

You will need:

* A Databricks workspace with **Unity Catalog** enabled. Legacy `hive_metastore` tables also work; see [Legacy Hive metastore tables](#legacy-hive-metastore-tables).
* **Workspace admin** rights, to create a service principal and generate an OAuth secret for it.
* A **SQL warehouse** Cotool can use. We recommend a dedicated **serverless** SQL warehouse:
  * Serverless warehouses start in seconds. Classic and pro warehouses can take several minutes to start after auto-stop, and Cotool's first queries after an idle period may time out while they start.
  * A dedicated warehouse keeps Cotool's queries and their cost separate from your other workloads. Every Cotool query runs on your warehouse and consumes DBUs. Set an auto-stop interval that suits your budget.
* The **owner** of the catalogs or schemas holding your security data (or a metastore admin), to grant read access.

<Note>
  Cotool queries through a SQL warehouse, not an all-purpose cluster. If the
  HTTP path you have starts with `sql/protocolv1/o/`, it belongs to a cluster;
  create or pick a SQL warehouse instead.
</Note>

## Set up Databricks

<Steps>
  <Step title="Create a service principal">
    1. In your workspace, click your username in the top bar and open **Settings > Identity and access**.
    2. Next to **Service principals**, click **Manage**, then **Add service principal > Add new**.
    3. Enter a recognizable name (for example `cotool`) and click **Add**.
    4. Open the service principal and copy its **Application ID** (a UUID). This is the **Client ID** in Cotool.
  </Step>

  <Step title="Give it Databricks SQL access">
    On the service principal's **Configurations** tab, make sure **Databricks SQL access** is checked, and click **Update**.

    <Note>
      This entitlement is easy to miss. Without it, the service principal cannot use any SQL warehouse, even with **Can use** permission on one, and Cotool's connection check fails with a permission error.
    </Note>
  </Step>

  <Step title="Generate an OAuth secret">
    1. On the service principal, open the **Secrets** tab and click **Generate secret**.
    2. Choose a lifetime (up to 730 days) and click **Generate**.
    3. Copy the **Secret** value immediately. Databricks shows it only once.

    <Warning>
      On Azure Databricks, the secret must be generated here, in Databricks. A Microsoft Entra ID app registration's client secret does not work, even for an Entra ID-managed service principal.
    </Warning>

    Secrets expire. Before then, generate a new one and update it in Cotool under **Platform > Integrations > Databricks > Edit**.
  </Step>

  <Step title="Grant access to a SQL warehouse">
    1. Open **SQL Warehouses** and select the warehouse Cotool should use.
    2. Click **Permissions**, add the service principal, choose **Can use**, and click **Add**.
    3. On the **Connection details** tab, copy the **Server hostname**. Copy the **HTTP path** (for example `/sql/1.0/warehouses/1234567890abcdef`) only if you want to pin this warehouse.

    If you leave the HTTP path blank in Cotool, Cotool picks a warehouse the service principal can use: serverless first, then a running warehouse, then pro over classic. If that warehouse is later deleted, Cotool switches to another one. Pin a warehouse when you need Cotool's queries and their cost on a specific one.
  </Step>

  <Step title="Grant read access to your security data">
    In the SQL editor (or **Catalog Explorer > Permissions**), grant the service principal read access to the catalog that holds your security data. In `GRANT` statements, refer to a service principal by its Application ID:

    ```sql theme={null}
    -- Replace `security` with your catalog and the UUID with the Application ID from step 1.
    GRANT USE CATALOG ON CATALOG security TO `11111111-2222-3333-4444-555555555555`;
    GRANT USE SCHEMA  ON CATALOG security TO `11111111-2222-3333-4444-555555555555`;
    GRANT SELECT      ON CATALOG security TO `11111111-2222-3333-4444-555555555555`;
    ```

    Privileges granted on a catalog apply to every schema and table in it, including ones created later. To limit Cotool to specific data, grant `USE SCHEMA` and `SELECT` on individual schemas or tables instead. `USE CATALOG` on the parent catalog is always required.

    Do not grant `MODIFY`, `ALL PRIVILEGES`, or ownership. Cotool only needs to read.
  </Step>
</Steps>

## Connect in Cotool

1. In Cotool, go to **Platform > Integrations > Databricks** and click **Connect**.
2. Enter the values you collected:

| Cotool field | Where it comes from |
| - | - |
| Workspace host | **Server hostname** from the warehouse's Connection details (step 4), or the workspace URL from your browser |
| SQL warehouse HTTP path | Optional. **HTTP path** from the warehouse's Connection details (step 4), to pin that warehouse |
| Authentication method | **OAuth service principal** (recommended) |
| Service principal client ID | The service principal's **Application ID** (step 1) |
| Service principal OAuth secret | The **Secret** value generated in step 3 |

3. Click **Connect**. Cotool exchanges the credentials for a token and reads (or picks) the SQL warehouse to confirm both work before saving. If the warehouse is not serverless, the connect form shows a warning, because its first query after auto-stop waits for it to start. If the check fails, the error explains which part to fix (see [Troubleshooting](#troubleshooting)).

Workspace hosts look like this:

| Cloud | Example host |
| - | - |
| AWS | `dbc-1234abcd-5678.cloud.databricks.com` |
| Azure | `adb-1234567890123456.7.azuredatabricks.net` |
| GCP | `1234567890123456.7.gcp.databricks.com` |

Cotool only sends credentials to Databricks-operated domains, including the GovCloud and Azure Government and China variants. Custom vanity domains are not supported; use the workspace's Databricks hostname.

### Personal access token (fallback)

If your organization cannot create service principals, choose **Personal access token** and paste a token instead of the client ID and secret. Generate it for a dedicated, least-privilege identity rather than a person's account, because Cotool can read whatever that identity can. Tokens expire, and workspace admins can disable them; when that happens Cotool's queries start failing with HTTP 401 until you paste a new token.

## Network access

If your workspace restricts inbound access, Cotool must be able to reach it:

* **IP access lists** — allow Cotool's egress IP addresses, listed on the Databricks connect form in Cotool (**Platform > Integrations > Databricks > Connect**).
* **Private Link with public access disabled** — Cotool cannot reach the workspace. Keep a public front-end endpoint restricted by IP access list instead.

Blocked connections show as `HTTP 403` (IP access list) or a connection failure before any response (Private Link or firewall).

## How Cotool queries Databricks

* **Discovery is incremental.** Agents list catalogs, then schemas, then tables, and describe only the tables they need. Discovery uses the Unity Catalog metadata APIs, which do not wake the SQL warehouse. Sampling rows does use the warehouse.
* **Results are capped.** Queries return 500 rows by default and at most 10,000, with a 4 MiB result cap. When more rows match, the response says `truncated: true`, and the agent narrows or aggregates the query.
* **Queries time out and are cancelled.** Each statement has a timeout (300 seconds by default, at most 30 minutes). Cotool cancels statements that exceed it, so abandoned queries do not keep running on your warehouse.
* **Queries are attributable.** Cotool's statements appear in Databricks **Query History** under the service principal, so you can audit and cost them.

Agents query fastest when security tables are laid out for time-bounded reads:

* Partition or liquid-cluster large log tables on an event date or timestamp column. Cotool shows agents these columns and tells them to filter on them first.
* Add table and column comments describing each source (for example, "AWS CloudTrail management events"). Agents see comments during discovery and use them to pick the right table.
* Use one table per log source rather than one mixed table, so agents do not have to filter by source type.

## Environment map

After you connect, and weekly after that, Cotool maps the workspace's Unity Catalog tables: each table's columns, partition and clustering columns, a few sample rows, and a summary of what it holds (for example, which log source it is). Agents start from this map instead of rediscovering the catalog, and detection planning uses it to know which log sources you have. View it on the Databricks integration page in Cotool.

* The map covers the tables the service principal can see, up to 300. It skips the `system`, `samples`, and `hive_metastore` catalogs and `information_schema` schemas; agents can still query those directly.
* Listing tables uses the Unity Catalog APIs and does not wake the warehouse. Sampling runs one `SELECT * ... LIMIT 3` per table, so mapping briefly starts the warehouse. If the warehouse cannot start, Cotool maps tables without sample rows.
* Tables created between runs are not in the map until the next run, but agents can still find them with discovery.

## Legacy Hive metastore tables

Tables in the workspace-local `hive_metastore` catalog are not served by the Unity Catalog APIs, so Cotool discovers them with `SHOW` and `DESCRIBE` statements on the SQL warehouse. This works, with these differences:

* Table-name search across the catalog is not available; agents browse schema by schema.
* Discovery wakes the SQL warehouse.
* Access is controlled by legacy table access control rather than Unity Catalog grants. Grant the service principal `USAGE` and `SELECT` on the relevant databases, or upgrade the tables to Unity Catalog.

## Troubleshooting

| Error mentions | What to fix |
| - | - |
| `invalid_client`, or "Databricks rejected the OAuth client credentials" | The client ID or secret is wrong, expired, or an Entra ID secret. Generate a new secret on the service principal's **Secrets** tab (step 3). |
| "is not a Databricks workspace host" | Enter the workspace's Databricks hostname (see the table above), not a vanity domain or a non-Databricks URL. |
| "belongs to an all-purpose cluster" | The HTTP path is for a cluster. Use a SQL warehouse's HTTP path (step 4). |
| `HTTP 403` / `PERMISSION_DENIED` on connect | The service principal lacks **Can use** on the warehouse (step 4) or the **Databricks SQL access** entitlement (step 2), or an IP access list blocks Cotool. |
| `HTTP 404` on connect | The HTTP path's warehouse ID does not exist in this workspace. Copy the HTTP path again from the warehouse's Connection details. |
| "No SQL warehouses are visible to this identity" | The HTTP path was left blank and the service principal has no warehouse permission. Grant it **Can use** on a SQL warehouse (step 4), preferably serverless. |
| `TABLE_OR_VIEW_NOT_FOUND`, or catalogs and tables missing from discovery | The service principal lacks `USE CATALOG`, `USE SCHEMA`, or `SELECT` on that data (step 5). Unity Catalog hides objects an identity cannot access. |
| "did not finish within N seconds" and "warehouse is STARTING" | A classic or pro warehouse was starting after auto-stop. Retry in a few minutes, or switch Cotool to a serverless warehouse. |
| "did not finish within N seconds" on a running warehouse | The query scanned too much data. Filter on the table's partition or clustering column, or give the warehouse more capacity. |
| Connection failure before any response | The workspace is not reachable from Cotool. See [Network access](#network-access). |

## Local development

Populate `cogent-backend/.env` with `DATABRICKS_HOST`, `DATABRICKS_HTTP_PATH`, `DATABRICKS_CLIENT_ID`, and `DATABRICKS_CLIENT_SECRET`, then run `npm run bootstrap-tools <API_KEY>` from the repo root.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.