The short answer
Combine sources. None sees everything, and each covers the other’s blind spot: identity shows what was authorized with the company account, the network shows where traffic goes, the endpoint shows what runs on the device, even away from the office, and finance shows what someone paid for. Asking people helps you understand the reason, but it does not work as a method: a good share of use is not declared, not out of bad faith, but because nobody thinks they need to say so.
This guide describes each source, what it sees and what it lets through, and proposes a seven-day plan to reach a first real inventory. If you are still at “what is this?”, start with What is Shadow AI.
69%
of organizations already suspect or have evidence that employees use prohibited public generative tools. Suspecting is not inventorying: the difference between the two is exactly what technical discovery resolves.
Gartner, Previsão sobre incidentes de Shadow AI até 2030 (Infosecurity Magazine)
The five data sources, side by side
| Source | What it sees well | What tends to slip through | Effort |
|---|---|---|---|
| Identity (single sign-on, app permissions) | Apps authorized with the company account and who authorized them. | Personal accounts and tools with no corporate login. | Low: data IT already has. |
| Network (DNS, proxy, firewall, SSE) | Corporate-network traffic to AI services, in volume. | Use off the network and on mobile data; weak attribution to the person. | Medium: access to and format of the logs. |
| Endpoint (agent) | What runs and what is accessed on the device, on any network, with attribution. | Devices without the agent, such as personal phones. | Medium: installation and rollout. |
| Browser (extension) | Browser use, in detail, if the extension inspects content. | Desktop apps and browsers without the extension; mind privacy. | Medium: rollout and policy. |
| Finance (invoices and expense claims) | Paid tools, including on an employee's card. | Everything free, which can be the largest share. | Low: data finance already has. |
1. Identity: single sign-on and app permissions
The company’s identity provider records which apps were authorized to access the account, and by whom. It is the cheapest source, because the data already exists, and it finds one specific kind of use: the AI app that asked for access to email, calendar or files with the corporate login. It can also reveal high risk, because those permissions can be broad.
What slips through: anything the person uses with a personal account, and tools that do not ask for a login.
2. Network: DNS, proxy, firewall and SSE
Network records show which services the traffic went to, in what volume and at what times. It is the best source for sizing use and for finding tools nobody remembered to declare, without installing anything on the devices. Formats vary across next-generation firewalls, secure access service edge (SSE) services and event platforms (SIEM), and the real work is in turning addresses into tools: one product uses several domains, and one domain can host several products.
What slips through: remote work outside the VPN, personal devices and encrypted traffic where only the destination is visible. Attribution to the person is also usually weak.
3. Endpoint: an agent on the device
A lightweight agent installed on workstations records what was accessed and which application ran, on any network, and knows which device and which user were involved. It is the most complete source for the company’s fleet of computers, and the one that solves remote work.
What slips through: devices without the agent, such as personal phones and computers. And there is a decision that has to be made and communicated beforehand: what the agent collects. The difference between recording that a tool was used and capturing what was typed is large, and it is covered in the privacy section below.
4. Browser: an extension
An extension sees what happens inside the browser, which is where a large share of generative AI use takes place, and can get down to the detail of what is typed. That detail is exactly what makes it the most sensitive source: collecting content creates a new repository of sensitive data, with the risks that brings.
What slips through: desktop apps and browsers without the extension.
5. Finance: invoices and expense claims
Subscriptions to AI tools paid on an employee’s card and later reimbursed show up in expenses, and are the most direct clue to recurring, relevant use. The source costs almost nothing, because finance already has the data.
What slips through: all free use, which can be the largest share and the one with the least control.
What still slips through
- AI built into software already approved. Productivity suites, support platforms and editors have started offering AI features inside the same product. The original approval did not cover that, and the traffic leaves from the same domain as ever.
- Personal devices and accounts. No corporate source covers them fully; the answer is policy and training, not only technology.
- Code and API keys. Technical teams call models straight from code, with keys that go through no login. The source here is the repository and the cloud, not the network.
- Local models. A model running on the machine itself generates no outbound traffic. Only the agent sees it.
A seven-day plan
| Day | What to do | Result |
|---|---|---|
| 1 | Pick the sources the company already has and ask for access: identity, network and finance. | Access to the data, without installing anything. |
| 2 | Extract the last 30 to 90 days of network records and app permissions. | A raw list of destinations and apps. |
| 3 | Turn addresses into tools by comparing against a catalog of AI tools. | A list of tools, not of domains. |
| 4 | Attribute each tool to a department and separate one-off from recurring use. | A map of use by area. |
| 5 | Classify each tool by risk and note the plan in use, where there is one. | A classified inventory. |
| 6 | Talk to the owners of the heaviest-use areas, without an accusatory tone. | The reason behind each use. |
| 7 | Decide what stays allowed, restricted or prohibited, and write the rule. | A policy draft and an action list. |
How to read the result
The first inventory tends to bring three groups: tools the company already knew about and approved, tools nobody knew were in use and that are harmless in practice, and a small group, with concentrated use, that deserves immediate attention. It serves less to find culprits and more to set priorities.
Two distinctions help. The first is between one-off and recurring use: a single visit to a site is not the same as a whole department depending on a tool. The second is the kind of data likely to go in: the same service carries different risks in marketing and in legal. To classify, our public catalog shows, per tool, the vendor, the country and the risk with the reason. The next step, handling each deviation, is in the guide on the risks of AI in the workplace.
92%
of organizations with an AI-related breach did not have adequate access control for those tools, according to the 2026 report. Access control presupposes knowing which tools you are giving access to: the inventory comes first.
The conversation with the people who use it
What people tell you after the survey is worth more than the survey itself. A few questions work well: what problem were you trying to solve?, what data did you put in? and what was missing from the approved tool? The tone matters. A conversation that opens with punishment can push use to where nobody can see it, and the company loses exactly the visibility it has just gained. Faced with real, frequent use, it is usually more productive to offer an approved alternative than to simply ban.
Discovering without becoming surveillance
There is a difference in kind between recording that a tool was used, from which workstation and when, and capturing what was typed. The first gives you the inventory and the evidence. The second creates a new repository of sensitive data, with privacy, leakage and employment-relationship risks.
In any design, inform people in advance, define the purpose in writing, limit collection to what is necessary and validate the arrangement with legal and the data protection officer, in light of the LGPD. Tangerin AI does discovery through an agent on workstations or through firewall, SSE and SIEM logs, and records which tool was used, when and on which workstation, without collecting the content of conversations, prompts or files. If you want to measure where your company stands before any project, the free assessment at /diagnostico asks nine questions, with no signup. If you are comparing vendors, see the ten criteria.
Frequently asked questions
How do you find out which AI tools employees use?
By combining sources, because none sees everything: single sign-on and app permissions show what was authorized with the company account; network logs (DNS, proxy, firewall, SSE) show where traffic goes; an agent on the workstation shows what runs on the device, even off the network; and invoices and expense claims show what someone paid for. Then you need to compare what showed up against a catalog that says which tool each address represents and what risk it carries.
Can I find out without installing anything on the machines?
In part. Network and identity logs give a good view of corporate-network traffic and authorized apps without installing anything. What stays out is use off the network, on personal devices and with personal accounts, which only an agent or an extension on the device can see.
How can I tell if someone uses a personal ChatGPT account for work?
Through identity, you cannot: a personal account does not go through the company login. What you can see is access to the service, from the network or the device, and its volume and frequency. What the company should not do is try to guess the account: what matters is knowing the tool is in use and with what data.
What is AI built into software we already use?
Many already-approved tools, such as productivity suites, support platforms and editors, have started offering AI features inside the same product. The original approval did not cover that. It is worth reviewing the terms and the data handling of those functions, and asking the vendor whether they are on by default.
Is it permitted to monitor employees' AI use?
It depends on what is collected, the purpose and the transparency. Recording which tool was used, from which workstation and when is very different from capturing what was typed. In any case, inform people, define the purpose in writing and validate the design with legal and the data protection officer before starting. This guide is not legal advice.
How often should the survey be repeated?
Continuously. The AI market changes every week: new tools appear, older ones change terms and some disappear. A list built once a year is born old. The ideal is an inventory that updates itself and warns when something new appears.
Sources
- IBM / Ponemon Institute — Cost of a Data Breach Report 2026 (2026). https://www.ibm.com/reports/data-breach
- Gartner — Previsão sobre incidentes de Shadow AI até 2030 (2025). via Infosecurity Magazine. https://www.infosecurity-magazine.com/news/gartner-40-firms-hit-shadow-ai/
- McKinsey & Company — The State of AI: Global Survey (2025). https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai
- Presidência da República — Lei n.º 13.709/2018 (LGPD) (2018). https://www.planalto.gov.br/ccivil_03/_ato2015-2018/2018/lei/l13709.htm