How a Workday career site loads its jobs
Many large employers run their job search on Workday. You can tell by the address: https://<tenant>.wd<N>.myworkdayjobs.com/<site>. The page you see is an app. It gets its job list from a JSON endpoint on the same host, with a POST request to:
https://<tenant>.wd<N>.myworkdayjobs.com/wday/cxs/<tenant>/<site>/jobs
That endpoint needs no login and no key. It is public because the career site itself calls it for every visitor. This guide reads it the same way, slowly, and only after checking the host's robots.txt.
Find the tenant, data centre and site
Open the company's careers page and click through to its job search. Mastercard's careers page, for example, links to:
https://mastercard.wd1.myworkdayjobs.com/CorporateCareers
That address has the three parts you need:
| Part | In the example | Notes |
|---|---|---|
| Tenant | mastercard | The first label of the host. It appears twice in the API path. |
| Data centre | wd1 | Varies by company: Mastercard uses wd1, NVIDIA wd5, Salesforce wd12. Copy it, do not guess it. |
| Site | CorporateCareers | The first path segment after an optional language part such as en-US. Copy it as written. |
So the job list of that site is at https://mastercard.wd1.myworkdayjobs.com/wday/cxs/mastercard/CorporateCareers/jobs. A company can have several sites (campus, a subsidiary); each one is a separate list. A wrong site name answers HTTP 404. A web search for site:myworkdayjobs.com plus the company name lists a company's sites.
This guide covers myworkdayjobs.com addresses. Some employers use a different Workday address style (wd5.myworkdaysite.com/...) or embed the search on their own domain; Tinlark has not tested those here.
Check robots.txt first
Each career-site host publishes a robots.txt. Read it before you request anything else, and skip the site if it disallows the path you want.
curl -s https://mastercard.wd1.myworkdayjobs.com/robots.txtSitemap: https://mastercard.wd1.myworkdayjobs.com/CorporateCareers/siteMap.xml
Sitemap: https://mastercard.wd1.myworkdayjobs.com/Campus/siteMap.xml
Sitemap: https://mastercard.wd1.myworkdayjobs.com/ContractorConversion/siteMap.xml
Sitemap: https://mastercard.wd1.myworkdayjobs.com/RYC/siteMap.xml
User-agent: *
Allow: /CorporateCareers/
Allow: /Campus/
Allow: /ContractorConversion/
Allow: /RYC/
Disallow: /CUOReqSite/
Disallow: /CampusApplyOnly/
Disallow: /Public_Posting_Site/
Disallow: /refreshFacet/Nothing here disallows /wday/cxs/, so the job list may be read. Other tenants publish other rules; the script below checks them with Python's urllib.robotparser.
Read the first page
Send a JSON body with limit, offset, an empty searchText and no filters:
curl -s -X POST "https://mastercard.wd1.myworkdayjobs.com/wday/cxs/mastercard/CorporateCareers/jobs" \
-H "Content-Type: application/json" -H "Accept: application/json" \
-d '{"appliedFacets":{},"limit":2,"offset":0,"searchText":""}' | python3 -m json.tool{
"total": 1037,
"jobPostings": [
{
"title": "Manager, Software Engineering",
"externalPath": "/job/Ramat-Gan-Israel/Manager--Software-Engineering_R-279658",
"locationsText": "Ramat-Gan, Israel",
"postedOn": "Posted Today",
"bulletFields": [
"R-279658"
]
},
{
"title": "Manager, Accounting",
"externalPath": "/job/Pune-India/Manager--Accounting_R-291473",
"locationsText": "Pune, India",
"postedOn": "Posted Today",
"bulletFields": [
"R-291473"
]
}
],
"facets": [
{
"facetParameter": "jobFamilyGroup",
"descriptor": "Job Category",
"values": [
{
"descriptor": "Engineering",
"id": "189119ebe266100103737c3d6a6e0000",
"count": 260
},Each posting has a title, a location text, a coarse posting age and bulletFields, whose first entry is usually the requisition id. The job's public page is the site address plus externalPath: https://mastercard.wd1.myworkdayjobs.com/CorporateCareers/job/Ramat-Gan-Israel/Manager--Software-Engineering_R-279658.
Page through the list
Three rules, all checked today:
- 20 jobs per request at most.
"limit": 50answers HTTP 400 with{"errorCode":"HTTP_400",...}. Use 20 and raiseoffsetby 20. totalis only on the first page. Later pages report"total": 0, so keep the number from page one.- Go slowly. One request a second is plenty: 1,037 jobs are 52 requests, and the script below exported them in about two minutes.
The 2,000-job ceiling
For large employers the list stops at 2,000. NVIDIA's career site reports "total": 2000, and asking for an offset of 2,000 or more returns the first page again:
for offset in 0 1980 2000; do
curl -s -X POST "https://nvidia.wd5.myworkdayjobs.com/wday/cxs/nvidia/NVIDIAExternalCareerSite/jobs" \
-H "Content-Type: application/json" \
-d "{\"appliedFacets\":{},\"limit\":20,\"offset\":$offset,\"searchText\":\"\"}" |
python3 -c "import json,sys; d=json.load(sys.stdin); print($offset, d['total'], [p['bulletFields'][0] for p in d['jobPostings'][:2]])"
sleep 1
done0 2000 ['JR2016289', 'JR2018381']
1980 0 ['JR2020205', 'JR2020393']
2000 2000 ['JR2016289', 'JR2018381']So a loop that trusts total stops at 2,000 jobs, and a loop that ignores it collects duplicates forever. Stop at offset >= min(total, 2000), and de-duplicate by requisition id.
Getting past 2,000: split by facet
The facets in the first answer count jobs per category, and those counts are not capped. Each filtered list has its own 2,000 window:
curl -s -X POST "https://nvidia.wd5.myworkdayjobs.com/wday/cxs/nvidia/NVIDIAExternalCareerSite/jobs" \
-H "Content-Type: application/json" \
-d '{"appliedFacets":{},"limit":1,"offset":0,"searchText":""}' |
python3 -c "
import json, sys
d = json.load(sys.stdin)
f = next(f for f in d['facets'] if f['facetParameter'] == 'jobFamilyGroup')
for v in f['values'][:4]:
print(v['count'], v['descriptor'], v['id'])
print('sum of all', len(f['values']), 'groups:', sum(v['count'] for v in f['values']))"1736 Engineering 0c40f6bd1d8f10ae43ffaefd46dc7e78
328 Sales 0c40f6bd1d8f10ae43ffcac5bbec7e90
121 Univ Employment 0c40f6bd1d8f10ae43ffda1e8d447e94
116 Operations 0c40f6bd1d8f10ae43ffc3fc7d8c7e8a
sum of all 14 groups: 2674The site lists 2,674 jobs, not 2,000. Put a category id into appliedFacets and page through that list instead:
curl -s -X POST "https://nvidia.wd5.myworkdayjobs.com/wday/cxs/nvidia/NVIDIAExternalCareerSite/jobs" \
-H "Content-Type: application/json" \
-d '{"appliedFacets":{"jobFamilyGroup":["0c40f6bd1d8f10ae43ffaefd46dc7e78"]},"limit":20,"offset":0,"searchText":""}' |
python3 -c "import json,sys; print(json.load(sys.stdin)['total'])"1736Facet names differ between tenants (Mastercard's are jobFamilyGroup, workerSubType, timeType and locationMainGroup), so read the facet names from the first answer, and de-duplicate by requisition id in case a job appears in two lists. If one group still holds more than 2,000 jobs, combine it with a second facet.
Full locations, exact dates and descriptions
The list is short on detail. Counted over the jobs.json that the script further down saved from Mastercard's site today:
import json
jobs = json.load(open("jobs.json"))
print(len(jobs), "jobs")
print(sum((j["location"] or "").endswith("Locations") for j in jobs), 'say "N Locations" instead of a place')
print(sum(not j["location"] for j in jobs), "have no location text")
print(sum(j["posted"] == "Posted 30+ Days Ago" for j in jobs), 'say "Posted 30+ Days Ago"')1037 jobs
100 say "N Locations" instead of a place
3 have no location text
314 say "Posted 30+ Days Ago"Each job has a detail endpoint: the API path plus the job's externalPath, read with GET. It returns every location, the exact posting date, the time type and the description as HTML.
curl -s "https://mastercard.wd1.myworkdayjobs.com/wday/cxs/mastercard/CorporateCareers/job/Dublin-Ireland/Director--Product-Management--Developer-Experience_R-279947" \
-H "Accept: application/json" |
python3 -c "
import json, sys
info = json.load(sys.stdin)['jobPostingInfo']
for k in ('title', 'jobReqId', 'location', 'additionalLocations', 'postedOn', 'startDate', 'timeType'):
print(k, '=', info.get(k))
print('jobDescription =', len(info['jobDescription']), 'characters of HTML')"title = Director, Product Management, Developer Experience
jobReqId = R-279947
location = Dublin, Ireland
additionalLocations = ['Lisbon, Portugal', 'Warsaw, Poland (Plac Europejski 1)']
postedOn = Posted 2 Days Ago
startDate = 2026-10-02
timeType = Full time
jobDescription = 4422 characters of HTMLThat is one extra request per job, so filter the list first (by title, by facet) and fetch details only for the jobs you keep.
A complete script: every job to JSON and CSV
Standard library only. It parses the careers URL, checks robots.txt, pages at one request a second, stops at the total or at 2,000, and writes both files. Put your own contact in the User-Agent.
"""Export every open job of one public Workday career site to jobs.json and jobs.csv.
Usage: python3 workday_jobs.py https://mastercard.wd1.myworkdayjobs.com/en-US/CorporateCareers"""
import csv, json, re, sys, time, urllib.request, urllib.robotparser
UA = "my-job-export/1.0 (contact: [email protected])" # say who you are
def parse(url):
m = re.match(r"https://([a-z0-9-]+)\.(wd\d+)\.myworkdayjobs\.com/(?:[a-z]{2}-[A-Z]{2}/)?([A-Za-z0-9_.-]+)", url)
if not m:
sys.exit("Not a myworkdayjobs.com career-site URL")
return m.groups() # tenant, data centre, site
tenant, wd, site = parse(sys.argv[1])
host = f"https://{tenant}.{wd}.myworkdayjobs.com"
api = f"{host}/wday/cxs/{tenant}/{site}/jobs"
robots = urllib.robotparser.RobotFileParser(f"{host}/robots.txt")
robots.read()
if not robots.can_fetch(UA, api):
sys.exit(f"robots.txt of {host} does not allow {api}")
jobs, seen, offset, total = [], set(), 0, None
while True:
body = json.dumps({"appliedFacets": {}, "limit": 20, "offset": offset, "searchText": ""}).encode()
req = urllib.request.Request(api, data=body, headers={
"Content-Type": "application/json", "Accept": "application/json", "User-Agent": UA})
with urllib.request.urlopen(req, timeout=30) as r:
page = json.load(r)
if total is None:
total = page["total"] # only the first page reports the total
posts = page.get("jobPostings") or []
for p in posts:
if p.get("externalPath") in seen:
continue
seen.add(p.get("externalPath"))
jobs.append({
"title": p.get("title"),
"jobId": (p.get("bulletFields") or [""])[0],
"location": p.get("locationsText"),
"posted": p.get("postedOn"),
"url": host + "/" + site + p.get("externalPath", ""),
})
offset += 20
if not posts or offset >= min(total, 2000): # the list stops at 2,000
break
time.sleep(1) # one request a second
json.dump(jobs, open("jobs.json", "w"), indent=2)
with open("jobs.csv", "w", newline="") as f:
w = csv.DictWriter(f, fieldnames=list(jobs[0]))
w.writeheader()
w.writerows(jobs)
print(f"{site}: {total} listed, {len(jobs)} exported")python3 workday_jobs.py https://mastercard.wd1.myworkdayjobs.com/en-US/CorporateCareers
head -3 jobs.csvCorporateCareers: 1037 listed, 1037 exported
title,jobId,location,posted,url
"Manager, Software Engineering",R-279658,"Ramat-Gan, Israel",Posted Today,https://mastercard.wd1.myworkdayjobs.com/CorporateCareers/job/Ramat-Gan-Israel/Manager--Software-Engineering_R-279658
"Manager, Accounting",R-291473,"Pune, India",Posted Today,https://mastercard.wd1.myworkdayjobs.com/CorporateCareers/job/Pune-India/Manager--Accounting_R-291473The script does not split by facet, so on sites above 2,000 jobs it saves the first 2,000; add the facet loop from above if you need them all.
What Workday's public list does not give you
- No department or team. The job category facet is the closest thing, and only as a filter.
- No structured pay. Where an employer publishes a range, it is inside the description text.
- No company name. You know the tenant (
mastercard), not a display name. - Coarse dates ("Posted 30+ Days Ago") unless you read each job's detail.
- Some sites do not answer scripted requests at all. Treat that as a no and move on.
When a hosted tool is easier
One career site, once, is a script. It gets harder when you watch thirty companies every morning and want only what is new: you need a schedule, a place to keep yesterday's ids, retries on HTTP 429 and 5xx, and the same columns for every company. That is what the Workday Jobs Scraper on the Apify Store does with the endpoint above. It returns one uniform row per job, reads up to 2,000 jobs per career site per run (it does not split by facet), and its new-only mode returns only jobs that were not there on the previous run. Run it on Apify's scheduler, or call it from the Apify API.
Tinlark is not affiliated with Workday, Inc., Mastercard, NVIDIA or Salesforce. Their names appear only to point at their public career sites. The postings belong to the employers: check their terms and the laws that apply to you before you republish them.