Website Technology Detector
In the Apify Store
The technology stack of every site in your list: CMS, shop platform, framework, analytics, CDN and hosting, with the evidence for each. One row per site, no API key.
Paste a list of websites and get the technology stack of each one: CMS, shop platform, JavaScript framework, analytics and tag managers, CDN and hosting, payment and chat widgets, cookie consent tools, fonts, maps and video players. Every technology comes with its category, the version when the site shows one, a confidence level and the evidence found (a header, cookie, meta tag, script URL or page marker). The Actor reads each site's home page once, checks the site's robots.txt first, identifies itself as TinlarkBot and needs no API key.
What you get
- 500+ technologies. The list covers 535 technologies in 35 categories, with fingerprints Tinlark wrote itself from vendors' public documentation and from reading real pages.
- Evidence and confidence on every technology.
highwhen a response header, cookie, meta generator tag or a script from the vendor's own host shows it;mediumfor a single page-level signal;lowwhen it is implied by another technology (PHP when WordPress is found). - Find sites that use a technology. Run a prospect or competitor list and keep the rows whose
technologyNamescontain Shopify, WooCommerce, HubSpot, Klaviyo or any other name you look for. - Filter by category. Set
categoriesto keep only, for example, E-commerce and Analytics. An empty list keeps everything. - Versions where the site shows them. A generator tag, a Server header or a file name with a version. Most technologies publish none, so
versionis oftennull. - Up to 5,000 sites per run. A bare domain is tried as
https://domain, thenhttps://www.domain, thenhttp://domain. Give a full URL to check a page other than the home page. - Polite by design. robots.txt is always honoured, requests to one host are at least one second apart, and the Actor uses no proxies and does not run JavaScript.
Real output
{
"url": "https://wordpress.org",
"finalUrl": "https://wordpress.org/",
"domain": "wordpress.org",
"status": "ok",
"httpStatus": 200,
"technologyCount": 9,
"technologies": [
{
"name": "WordPress",
"category": "CMS",
"version": "7.2",
"confidence": "high",
"evidence": [
"header:link",
"meta:generator",
"link:/wp-content/"
]
},
{
"name": "Google Tag Manager",
"category": "Tag managers",
"version": null,
"confidence": "high",
"evidence": [
"iframe:googletagmanager.com/ns.html",
"html:GTM-P24PF4B"
]
},
{
"name": "Nginx",
"category": "Web servers",
"version": null,
"confidence": "high",
"evidence": [
"header:server"
]
},
{
"name": "PHP",
"category": "Backend",
"version": null,
"confidence": "low",
"evidence": [
"implied by WordPress"
]
}
],
"technologyNames": [
"Jetpack Stats",
"Jetpack Site Accelerator",
"WordPress",
"Google Fonts",
"Google Tag Manager",
"Nginx",
"Jetpack",
"Gutenberg blocks",
"PHP"
],
"categories": {
"Analytics": [
"Jetpack Stats"
],
"Backend": [
"PHP"
],
"CDN": [
"Jetpack Site Accelerator"
],
"CMS": [
"WordPress"
],
"Fonts": [
"Google Fonts"
],
"Tag managers": [
"Google Tag Manager"
],
"Web servers": [
"Nginx"
],
"WordPress": [
"Jetpack",
"Gutenberg blocks"
]
},
"server": "nginx",
"poweredBy": null,
"generator": "WordPress 7.2-alpha-64071",
"checkedAt": "2026-10-03T14:52:21Z",
"error": null
}
Example input
Paste this into the input form, or send it through the API.
{
"urls": [
"https://stripe.com",
"https://shopify.com",
"https://wordpress.org"
],
"categories": [],
"includeEvidence": true,
"includeVersions": true
}
Run it from code
One request to the Apify API starts a run and returns the dataset rows when it finishes. Use your own Apify API token. The Apify client libraries and Apify's MCP server work too.
API=https://api.apify.com/v2/acts/tinlark~website-technology-detector
curl -X POST "$API/run-sync-get-dataset-items" \
-H "Authorization: Bearer $APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"urls": ["stripe.com", "shopify.com"], "categories": ["E-commerce", "Analytics"]}'
Limits
- No JavaScript is executed. The Actor reads the HTML, headers and cookies the server sends. A site that builds its page in the browser shows only what is in the first HTML, and tools that a tag manager loads after the page starts are not visible.
- One page per site. Only the URL you give is fetched (a bare domain means its home page). Other pages of the site may use other tools.
- A missing technology means "not visible", not "not used". Many sites hide their server software, language and framework, so back-end and hosting facts are inferred from what the server reveals.
- Versions are rare.
versionisnullunless the site publishes one. - Some sites will not answer. In a run of 100 well-known sites, 97 were analysed, 1 disallowed TinlarkBot in robots.txt and 2 were unreachable. Sites behind bot protection often answer 403; the Actor reports that and does not use proxies or other means to get past it.
- Fingerprints are Tinlark's own. A run of 100 well-known sites found 177 different technologies. Fingerprints for rarer tools follow the vendors' documented embed code but may not yet have been seen on a live page, so a rare tool can be missed or, rarely, named wrongly.
- Pace. A mixed list of 100 sites took 58 seconds at 1,024 MB (the default) and 212 seconds at 256 MB. The default run Timeout is 2 hours; raise it for very long lists of slow sites.
Data source
The listed sites themselves: their public home pages, response headers and robots.txt, read only where robots.txt allows. The fingerprint list is written by Tinlark and ships with the Actor; the Actor calls no third-party API and sends the sites you list to no one else. You are responsible for how you use the output and for following the terms of the sites you list.
Not affiliated with any of the technologies or companies it detects.
Questions before you start?
Write to [email protected]. Each product page lists what the tool does, what it does not do, and its exact price.