Domain Tech Stack, Contacts & SEO Profiler
In the Apify Store
One profile per company domain: technology stack, published business contacts, social links, SEO meta, sitemap size and security headers.
Paste a list of company domains and get one structured row per domain. Optional DNS, email authentication (SPF, DMARC) and domain registration data. It reads each site's home page and a few standard pages, only where robots.txt allows.
What you get
- Technology stack. About 160 technologies (CMS, shop platform, framework, analytics, tag manager, CRM, chat, payments, CDN, server), each with a category, a version when the site shows one, a confidence level and the evidence found.
- Business contacts from the company's own pages. Role addresses such as
info@orsales@on the company's domain and phone numbers the site publishes. Business-only by default. - Social links the site itself links to: LinkedIn company page, X, Facebook, Instagram, YouTube, GitHub, TikTok.
- SEO meta. Title, description, canonical, Open Graph tags, H1 count, language, schema.org types, hreflang count,
noindexflag. - Sitemap size and security headers. URL count of the XML sitemap; HSTS, CSP, framing and other headers with a simple A to F grade.
- DNS and domain age (optional). Nameservers, mail provider, SPF and DMARC policy, registrar, registration and expiry dates.
Real output
{
"recordType": "profile",
"domain": "hubspot.com",
"finalUrl": "https://www.hubspot.com/",
"status": "ok",
"httpStatus": 200,
"robotsAllowed": true,
"pagesFetched": 4,
"title": "HubSpot | Software & Tools for your Business - Homepage",
"metaDescription": "HubSpot's customer platform includes all the marketing, sales, customer service, and CRM software you need to grow your business.",
"h1Count": 1,
"lang": "en",
"schemaTypes": ["AggregateRating", "Brand", "Organization", "Product", "WebSite"],
"technologies": [
{ "name": "Cloudflare", "category": "CDN", "version": null, "confidence": "high", "evidence": ["header:server", "header:cf-ray", "cookie"] },
{ "name": "HubSpot CMS", "category": "CMS", "version": null, "confidence": "high", "evidence": ["html:/hubfs/", "meta:generator"] }
],
"contacts": { "emails": [], "phones": ["18884827768"], "contactPageUrl": "https://offers.hubspot.com/contact-sales", "policy": "business-only" },
"socials": { "linkedinCompany": "https://www.linkedin.com/company/hubspot", "x": "https://x.com/HubSpot", "github": null },
"sitemap": { "url": "https://www.hubspot.com/sitemap.xml", "urlCount": 3066, "lastmod": "2026-10-02", "isIndex": false },
"securityHeaders": { "hsts": true, "csp": true, "xFrameOptions": "DENY", "score": 7, "grade": "A" },
"dns": { "nameservers": ["jerry.ns.cloudflare.com", "yolanda.ns.cloudflare.com"], "mxProvider": "Google Workspace", "dmarcPolicy": "reject" },
"rdap": { "registrar": "MarkMonitor Inc.", "createdAt": "2005-02-06T20:02:28Z", "expiresAt": "2027-02-06T20:02:28Z", "domainAgeDays": 7907 }
}
Example input
Paste this into the input form, or send it through the API.
{
"domains": ["notion.so", "linear.app", "figma.com", "vercel.com", "cohere.com"],
"modules": ["tech", "contacts", "socials", "meta"]
}
Run it from code
One request to the Apify API starts a run and returns the dataset rows when it finishes. Use your own Apify API token. The Apify client libraries and Apify's MCP server work too.
API=https://api.apify.com/v2/acts/tinlark~domain-intelligence-profiler
curl -X POST "$API/run-sync-get-dataset-items" \
-H "Authorization: Bearer $APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"domains": ["notion.so", "linear.app"], "modules": ["tech", "contacts", "socials", "meta"]}'
Limits
- Some sites will not answer. In a test of 98 well-known company domains, 82 returned a profile. The rest refused automated clients, disallowed crawlers or did not respond. It does not use proxies or other means to get past blocks.
- Contacts are best effort. In a test of 50 software company domains, a business email was found for 18, a phone number for 11 and a LinkedIn company link for 33.
- It does not crawl. It reads the home page, a few contact, imprint or about pages linked from it (three by default), robots.txt and the sitemap.
- No JavaScript. Detection reads the HTML the server sends, so tools loaded only by scripts or a tag manager can be missed.
- Pace. Requests to one site are at least one second apart. Plan about 70 minutes for 1,000 domains.
Data source
The listed sites themselves (public pages, headers, robots.txt and sitemaps, only where robots.txt allows), Cloudflare public DNS over HTTPS, and the registries' own RDAP servers. It identifies itself as TinlarkBot and never logs in, solves captchas or gets around blocks.
Not affiliated with any of the technologies or companies it detects.
Questions before you start?
Write to [email protected]. Each product page lists what the tool does, what it does not do, and its exact price.