# Database scan

{/* AUTO-SYNCED SOURCE: this page lives in apps/app/src/modules/db-leads/docs/ and is mirrored into the docs app by `bun sync:module-docs`. Edit it in the module, not in apps/docs. */}

<Lead>
The **database scan** is the core feature of DB-Leads: it works through every imported contact and its activity history for seller signals and files evidence-backed findings into the review inbox. The principle: traverse everything, analyze selectively, surface filtered results.
</Lead>

## Starting a scan

Start a run from the tool page via the **Datenbank-Scan** button. It requires the scan permission (`db-leads.scan`). The scan runs in the background. You keep working, and the tool page shows progress and stats live.

<Tip title="Import first, then scan">
The scan analyzes your contacts' **activities**: emails, calls, appointments, notes. The more complete the [onOffice import](/tools/db-leads/onoffice-import) (including activities), the more the scan has to read. If a sync is running, the scan starts right after it.
</Tip>

## The two stages

### Stage 1: Prefilter (deterministic, free)

The prefilter checks **every** contact, and no AI cost is incurred here:

<DefinitionList>
  <DefItem term="Time window">Only contacts whose last contact falls inside the time window (default 5 years, configurable 1 to 15) proceed to analysis. Never contacted? The creation date counts instead. Contacts outside the window are counted but not analyzed: no AI spend on dead records.</DefItem>
  <DefItem term="Exclusions">Contact types you excluded under **Settings → Qualification** are skipped by the scan.</DefItem>
  <DefItem term="Freshness check">Incremental: if a contact already has a briefing newer than its newest activity, there is nothing new to analyze, and it is skipped. That keeps every follow-up scan cheap.</DefItem>
  <DefItem term="Keyword boost">Activity texts are matched against your signal keywords (e.g. "Folgeauftrag"). Hits go first in the analysis order. Keywords prioritize, they never exclude: the AI also finds signals without a keyword.</DefItem>
</DefinitionList>

### Stage 2: AI deep analysis (per candidate)

For each candidate that passes the prefilter, the AI produces a structured briefing: seller signals, sales-readiness, summary, next step. **Every statement carries evidence**: a quote from a concrete activity with its date. A signal without evidence is discarded.

The analysis focuses on your **correspondence**: email bodies and call notes are read in full. That is where the passing remarks announcing a sale sit, and that is where nobody looks a second time in day-to-day work.

The same reading produces the **property leads**: which properties the person owns, where they are, how big they are, when they were built. They appear on the contact under **Ownership**, each with the sentence it came from. In this phase that is the only source for them, because before the mandate the property exists in no record. See [Ownership](/tools/db-leads/objekte).

## Where an indication comes from

For every signal the scan records what it rests on. It is shown on each finding so you can tell at a glance whether a call is worth it:

<DefinitionList>
  <DefItem term="From the correspondence">The person said or wrote it themselves: email, call note, appointment note, remark. It carries a date. This is the discovery, and a finding grows out of it.</DefItem>
  <DefItem term="CRM field">A field your office maintains. It frames the assessment: how this contact is handled, what has already been discussed, where they stand in the process. A valuation request means something different when the status field already says "sole mandate signed". That is exactly what the analysis uses your fields for.</DefItem>
  <DefItem term="Record data">Contact type, characteristic, origin of the record. Frames the picture like a CRM field does.</DefItem>
</DefinitionList>

**A finding needs at least one piece of evidence from the correspondence.** The reason: whatever sits only in your own fields is something you already know. The findings view is a worklist for what is new, not a second view of your record data. The briefing stays on the contact either way, it just does not enter the list.

**Age counts.** Remarks are the most valuable content in your database and often several years old. An indication from last year therefore weighs more than the same sentence from 2019. Old indications do not disappear, they land in the lowest band. For a CRM field the date is unknown: you know what the state is, but not since when.

## Seller score & bands

From the evidenced signals the tool computes a deterministic **seller score of 0 to 100**. No model decides it; the same signals always produce the same score:

<DefinitionList>
  <DefItem term="Hot (score 60 and up)">Strong, recent selling signals, for example a mentioned follow-up mandate plus a valuation request. Work through these first.</DefItem>
  <DefItem term="Warm (40 to 59)">Clear indications, not yet conclusive: an expired sole mandate or a single life event.</DefItem>
  <DefItem term="Watch (20 to 39)">A single evidenced signal, for example a passing mention of a follow-up mandate. It reaches the inbox because a glance costs little and a missed listing costs a lot.</DefItem>
  <DefItem term="No finding (below 20)">Too weak. A finding is only created from **score 20** upwards, and only with evidence from the correspondence.</DefItem>
</DefinitionList>

A single, clearly evidenced core signal from the correspondence is enough for a finding. That is deliberate: the inbox costs you a glance, a missed listing costs you a commission.

## Adjusting sensitivity

You decide from which point a contact appears as a finding: under Settings, Database scan, in the **Sensitivity** section. Lower means more findings and more work, higher means fewer and clearer cases. A well-maintained database tolerates a higher setting than one grown over twenty years.

Next to the slider you see what the setting would have produced on the last run, for example "this setting would have produced 47 findings instead of 12". It costs no new scan: the calculation runs on contacts that have already been assessed.

<Callout tone="note" title="A guide, not a promise">
The preview says what would have been. The next run sees new activity and may therefore differ.
</Callout>

## Analysis: what became of the findings

The same tab shows what became of your findings: how many you accepted, how many of those turned into a success, and which signals actually carry weight for you. If "life event" leads to an appointment in three out of four cases and "reaction" in none, you know what to look for.

**You define what counts as a success.** Under Settings, Pipeline, you can mark any phase as a success. For one office that is "appointment scheduled", for the next only "mandate signed". When a contact reaches such a phase after an accepted finding, this is recorded on the finding automatically. Nothing to enter by hand.

Below that you see why you dismissed findings. If one reason accumulates, that is rarely on you: usually the scan suggests too broadly. When one reason explains almost every fourth dismissed finding, the tool suggests a stricter sensitivity. Suggested, never changed automatically.

<Callout tone="note" title="Rates only from eight cases">
A rate based on three cases jumps between 0 and 100 percent depending on how one of them turns out. That is why it says "not enough yet" until enough findings have been decided. A finding from yesterday that is still running does not count as a failure.
</Callout>

## Learning from your results

If you like, the tool adjusts the scoring itself: signals that lead to a mandate more often in your office come to weigh more over time, signals with little return weigh less. The switch sits under **Settings · Database scan**, and it is **off** until you turn it on.

How it works:

<DefinitionList>
  <DefItem term="When">After every scan, and only once enough findings have been decided. Nothing is derived from three cases.</DefItem>
  <DefItem term="Measured against">Your own success rate, not a fixed value. A signal at 20 percent is good when your office sits at 8 percent.</DefItem>
  <DefItem term="How much">One small step per run at most. A single outlier does not rewrite your scoring.</DefItem>
  <DefItem term="The limits">No signal drops below half its starting value, none rises above one and a half times it.</DefItem>
</DefinitionList>

<Callout tone="note" title="Only the weights move">
No new attribute is added and none is removed. Which attributes may be scored at all is a reviewed list, not a question of statistics. What shifts is only how heavily an already reviewed attribute weighs.
</Callout>

In the analysis, **How the scoring adjusted** lists every change with date, direction and reason, including the numbers it followed from. One click on **Reset** puts all weights back to their starting values; the history stays, so what applied before remains traceable.

Turning the switch off freezes the weights where they are. Switching off and resetting are two different things.

## The background run

While scanning, the tool page shows the live status with four numbers: **checked**, **analyzed**, **findings**, and **outside the time window**, for example "3,482 checked · 214 analyzed · 12 findings · 1,240 outside the time window (5 years)".

<DefinitionList>
  <DefItem term="Resumable">If the run is interrupted, it picks up from the saved position. Nothing is analyzed twice.</DefItem>
  <DefItem term="Always runs to completion">A started scan never stops mid-run: the scan's AI analyses are included in the DB-Leads add-on and consume no credits. By default ONE run works through the entire database in stages until every candidate is analyzed (scales to 15,000+ contacts). Optionally set an **AI limit per run** in Settings: the run then finishes cleanly at the limit, e.g. "500 of 890 candidates analyzed", and the next scan continues exactly there (the freshness check skips what was already analyzed).</DefItem>
  <DefItem term="Cancel">Possible at any time. Findings already written are kept.</DefItem>
  <DefItem term="Full scan as a batch">A full scan submits all contacts to the AI provider at once instead of working through them one by one. The result arrives within 24 hours, you get a notification, and the run needs nothing from you until then. The scan card tells you it has been submitted. The regular scan (new material only) still runs immediately.</DefItem>
  <DefItem term="Nothing responding">If there is no progress for five minutes, the scan card says so. Such runs are closed automatically, a new scan continues from there, and findings already written are kept.</DefItem>
</DefinitionList>

## What is included?

The database scan is part of the DB-Leads add-on:

<DefinitionList>
  <DefItem term="Database scan (incremental)">Analyzes only contacts with new activity since the last analysis. Included in the add-on, start it any time.</DefItem>
  <DefItem term="Full scan">Re-analyzes every contact inside the time window. The add-on includes a fixed quota of full scans per month; the limit applies only at start, never mid-run.</DefItem>
  <DefItem term="Single briefings">A manually triggered AI briefing for one lead is billed against your credit balance as usual.</DefItem>
  <DefItem term="BYOK (Max plan and up)">Optionally the AI runs on your own provider key (Settings → REOS AI). You then pay the provider directly, no credits involved. Available from the Max plan.</DefItem>
</DefinitionList>

## Reviewing findings

Next to Board and List sits the **Findings** view, the scan's review inbox. Every finding shows the score, band, signals with evidence quotes, and a one-liner. Three actions:

<DefinitionList>
  <DefItem term="Accept">The lead moves into the recommended pipeline stage, and it is normal acquisition work from there. Requires the qualify permission (`db-leads.qualify`).</DefItem>
  <DefItem term="Dismiss">With an optional reason. Your reasons help improve the prefilter and analysis over time.</DefItem>
  <DefItem term="Snooze">Puts the finding aside. If a later scan adds new signals, it resurfaces.</DefItem>
</DefinitionList>

There is **at most one open finding per lead**: new scans update the open finding's score and signals instead of flooding the inbox.

### Working as a team: "Next finding"

Above the list sits **Nächster Fund** ("next finding"). One click takes the highest-scoring free finding, opens the contact, and reserves it for the lock time from **Settings → General → Queue**. While that runs, nobody else is offered this finding as their next one. Two people clicking at the same moment get different findings, never the same one.

Reserved does **not** mean locked. The finding stays visible in list, board, and search, and any colleague with the qualify permission can still accept, dismiss, or snooze it. The row then says who holds it and until when. The lock is an agreement, not a barrier.

Three things release a finding: a decision on it, the lock running out, and the setting `0` (findings are still handed out, nothing is reserved). Closing the tab releases nothing and needs to release nothing.

**Visible on the contact, too.** On the board and in the list, a reserved contact carries a small mark with the first name and the time the reservation runs until. If several people are working on the same contact (it can carry several findings), the mark says so. It disappears on its own once the time has passed. The same rule applies here: the contact opens and edits like any other.

## The time window (X years)

The time window under **Settings → Database scan** controls both what the scan analyzes **and** what the inbox shows:

- Only findings whose lead has a last contact **inside** the window are displayed. Change the window and the display adapts **immediately**: findings are shown or hidden, never deleted.
- Contacts outside the window are counted by the scan but not analyzed. You always see how much potential sits outside. Raise the window, and those contacts automatically become candidates on the next scan.

<Callout tone="note" title="AGG-compliant by design">
Seller signals come exclusively from the **person's own statements** in the activities, never from personal attributes such as age or marital status. "The house is getting too big for us" as a statement counts; a date of birth never does. The score only knows signal categories, and a human decides on every finding.
</Callout>

## Frequently asked questions

<Faq>
  <FaqItem q="Why wasn't a specific contact analyzed?">
    Three possible reasons: last contact outside the time window, excluded contact type, or the briefing is already newer than the newest activity (freshness check). The scan stats report the skipped contacts separately.
  </FaqItem>
  <FaqItem q="Am I missing signals outside the time window?">
    The scan deliberately doesn't analyze them: a contact last spoken to six years ago is rarely a useful lead. Keyword hits outside the window are still counted, though: you decide whether a larger window is worth it.
  </FaqItem>
</Faq>
