Export every open job from a Workday career site as JSON or CSV

How a public Workday career site lists its jobs, how to find the tenant and site in a careers URL, how paging works, the 2,000-job ceiling, and a short Python script that saves every job as JSON and CSV.

Last checked: . Every command and output below was run on that day; where output is trimmed, the page says so.

What you will get

  • The JSON endpoint behind every myworkdayjobs.com career site, and how to build its address from the careers URL.
  • Paging that works (20 jobs per request) and the limits you hit: no total after page one, and a list that stops at 2,000 jobs.
  • A way past 2,000 jobs by splitting the list by job category.
  • A Python script of about 50 lines, standard library only, that saves every job as jobs.json and jobs.csv.

How a Workday career site loads its jobs

Many large employers run their job search on Workday. You can tell by the address: https://<tenant>.wd<N>.myworkdayjobs.com/<site>. The page you see is an app. It gets its job list from a JSON endpoint on the same host, with a POST request to:

https://<tenant>.wd<N>.myworkdayjobs.com/wday/cxs/<tenant>/<site>/jobs

That endpoint needs no login and no key. It is public because the career site itself calls it for every visitor. This guide reads it the same way, slowly, and only after checking the host's robots.txt.

Find the tenant, data centre and site

Open the company's careers page and click through to its job search. Mastercard's careers page, for example, links to:

https://mastercard.wd1.myworkdayjobs.com/CorporateCareers

That address has the three parts you need:

PartIn the exampleNotes
TenantmastercardThe first label of the host. It appears twice in the API path.
Data centrewd1Varies by company: Mastercard uses wd1, NVIDIA wd5, Salesforce wd12. Copy it, do not guess it.
SiteCorporateCareersThe first path segment after an optional language part such as en-US. Copy it as written.

So the job list of that site is at https://mastercard.wd1.myworkdayjobs.com/wday/cxs/mastercard/CorporateCareers/jobs. A company can have several sites (campus, a subsidiary); each one is a separate list. A wrong site name answers HTTP 404. A web search for site:myworkdayjobs.com plus the company name lists a company's sites.

This guide covers myworkdayjobs.com addresses. Some employers use a different Workday address style (wd5.myworkdaysite.com/...) or embed the search on their own domain; Tinlark has not tested those here.

Check robots.txt first

Each career-site host publishes a robots.txt. Read it before you request anything else, and skip the site if it disallows the path you want.

Command
curl -s https://mastercard.wd1.myworkdayjobs.com/robots.txt
Output (complete)
Sitemap: https://mastercard.wd1.myworkdayjobs.com/CorporateCareers/siteMap.xml
Sitemap: https://mastercard.wd1.myworkdayjobs.com/Campus/siteMap.xml
Sitemap: https://mastercard.wd1.myworkdayjobs.com/ContractorConversion/siteMap.xml
Sitemap: https://mastercard.wd1.myworkdayjobs.com/RYC/siteMap.xml

User-agent: *
Allow: /CorporateCareers/
Allow: /Campus/
Allow: /ContractorConversion/
Allow: /RYC/
Disallow: /CUOReqSite/
Disallow: /CampusApplyOnly/
Disallow: /Public_Posting_Site/
Disallow: /refreshFacet/

Nothing here disallows /wday/cxs/, so the job list may be read. Other tenants publish other rules; the script below checks them with Python's urllib.robotparser.

Read the first page

Send a JSON body with limit, offset, an empty searchText and no filters:

Command
curl -s -X POST "https://mastercard.wd1.myworkdayjobs.com/wday/cxs/mastercard/CorporateCareers/jobs" \
  -H "Content-Type: application/json" -H "Accept: application/json" \
  -d '{"appliedFacets":{},"limit":2,"offset":0,"searchText":""}' | python3 -m json.tool
Output (trimmed after the first facet value; the full answer lists four facets)
{
    "total": 1037,
    "jobPostings": [
        {
            "title": "Manager, Software Engineering",
            "externalPath": "/job/Ramat-Gan-Israel/Manager--Software-Engineering_R-279658",
            "locationsText": "Ramat-Gan, Israel",
            "postedOn": "Posted Today",
            "bulletFields": [
                "R-279658"
            ]
        },
        {
            "title": "Manager, Accounting",
            "externalPath": "/job/Pune-India/Manager--Accounting_R-291473",
            "locationsText": "Pune, India",
            "postedOn": "Posted Today",
            "bulletFields": [
                "R-291473"
            ]
        }
    ],
    "facets": [
        {
            "facetParameter": "jobFamilyGroup",
            "descriptor": "Job Category",
            "values": [
                {
                    "descriptor": "Engineering",
                    "id": "189119ebe266100103737c3d6a6e0000",
                    "count": 260
                },

Each posting has a title, a location text, a coarse posting age and bulletFields, whose first entry is usually the requisition id. The job's public page is the site address plus externalPath: https://mastercard.wd1.myworkdayjobs.com/CorporateCareers/job/Ramat-Gan-Israel/Manager--Software-Engineering_R-279658.

Page through the list

Three rules, all checked today:

  • 20 jobs per request at most. "limit": 50 answers HTTP 400 with {"errorCode":"HTTP_400",...}. Use 20 and raise offset by 20.
  • total is only on the first page. Later pages report "total": 0, so keep the number from page one.
  • Go slowly. One request a second is plenty: 1,037 jobs are 52 requests, and the script below exported them in about two minutes.

The 2,000-job ceiling

For large employers the list stops at 2,000. NVIDIA's career site reports "total": 2000, and asking for an offset of 2,000 or more returns the first page again:

Command
for offset in 0 1980 2000; do
  curl -s -X POST "https://nvidia.wd5.myworkdayjobs.com/wday/cxs/nvidia/NVIDIAExternalCareerSite/jobs" \
    -H "Content-Type: application/json" \
    -d "{\"appliedFacets\":{},\"limit\":20,\"offset\":$offset,\"searchText\":\"\"}" |
  python3 -c "import json,sys; d=json.load(sys.stdin); print($offset, d['total'], [p['bulletFields'][0] for p in d['jobPostings'][:2]])"
  sleep 1
done
Output (offset, total, first two requisition ids)
0 2000 ['JR2016289', 'JR2018381']
1980 0 ['JR2020205', 'JR2020393']
2000 2000 ['JR2016289', 'JR2018381']

So a loop that trusts total stops at 2,000 jobs, and a loop that ignores it collects duplicates forever. Stop at offset >= min(total, 2000), and de-duplicate by requisition id.

Getting past 2,000: split by facet

The facets in the first answer count jobs per category, and those counts are not capped. Each filtered list has its own 2,000 window:

Command
curl -s -X POST "https://nvidia.wd5.myworkdayjobs.com/wday/cxs/nvidia/NVIDIAExternalCareerSite/jobs" \
  -H "Content-Type: application/json" \
  -d '{"appliedFacets":{},"limit":1,"offset":0,"searchText":""}' |
python3 -c "
import json, sys
d = json.load(sys.stdin)
f = next(f for f in d['facets'] if f['facetParameter'] == 'jobFamilyGroup')
for v in f['values'][:4]:
    print(v['count'], v['descriptor'], v['id'])
print('sum of all', len(f['values']), 'groups:', sum(v['count'] for v in f['values']))"
Output
1736 Engineering 0c40f6bd1d8f10ae43ffaefd46dc7e78
328 Sales 0c40f6bd1d8f10ae43ffcac5bbec7e90
121 Univ Employment 0c40f6bd1d8f10ae43ffda1e8d447e94
116 Operations 0c40f6bd1d8f10ae43ffc3fc7d8c7e8a
sum of all 14 groups: 2674

The site lists 2,674 jobs, not 2,000. Put a category id into appliedFacets and page through that list instead:

Command
curl -s -X POST "https://nvidia.wd5.myworkdayjobs.com/wday/cxs/nvidia/NVIDIAExternalCareerSite/jobs" \
  -H "Content-Type: application/json" \
  -d '{"appliedFacets":{"jobFamilyGroup":["0c40f6bd1d8f10ae43ffaefd46dc7e78"]},"limit":20,"offset":0,"searchText":""}' |
python3 -c "import json,sys; print(json.load(sys.stdin)['total'])"
Output
1736

Facet names differ between tenants (Mastercard's are jobFamilyGroup, workerSubType, timeType and locationMainGroup), so read the facet names from the first answer, and de-duplicate by requisition id in case a job appears in two lists. If one group still holds more than 2,000 jobs, combine it with a second facet.

Full locations, exact dates and descriptions

The list is short on detail. Counted over the jobs.json that the script further down saved from Mastercard's site today:

Command: python3 count.py
import json
jobs = json.load(open("jobs.json"))
print(len(jobs), "jobs")
print(sum((j["location"] or "").endswith("Locations") for j in jobs), 'say "N Locations" instead of a place')
print(sum(not j["location"] for j in jobs), "have no location text")
print(sum(j["posted"] == "Posted 30+ Days Ago" for j in jobs), 'say "Posted 30+ Days Ago"')
Output
1037 jobs
100 say "N Locations" instead of a place
3 have no location text
314 say "Posted 30+ Days Ago"

Each job has a detail endpoint: the API path plus the job's externalPath, read with GET. It returns every location, the exact posting date, the time type and the description as HTML.

Command
curl -s "https://mastercard.wd1.myworkdayjobs.com/wday/cxs/mastercard/CorporateCareers/job/Dublin-Ireland/Director--Product-Management--Developer-Experience_R-279947" \
  -H "Accept: application/json" |
python3 -c "
import json, sys
info = json.load(sys.stdin)['jobPostingInfo']
for k in ('title', 'jobReqId', 'location', 'additionalLocations', 'postedOn', 'startDate', 'timeType'):
    print(k, '=', info.get(k))
print('jobDescription =', len(info['jobDescription']), 'characters of HTML')"
Output
title = Director, Product Management, Developer Experience
jobReqId = R-279947
location = Dublin, Ireland
additionalLocations = ['Lisbon, Portugal', 'Warsaw, Poland (Plac Europejski 1)']
postedOn = Posted 2 Days Ago
startDate = 2026-10-02
timeType = Full time
jobDescription = 4422 characters of HTML

That is one extra request per job, so filter the list first (by title, by facet) and fetch details only for the jobs you keep.

A complete script: every job to JSON and CSV

Standard library only. It parses the careers URL, checks robots.txt, pages at one request a second, stops at the total or at 2,000, and writes both files. Put your own contact in the User-Agent.

workday_jobs.py
"""Export every open job of one public Workday career site to jobs.json and jobs.csv.
Usage: python3 workday_jobs.py https://mastercard.wd1.myworkdayjobs.com/en-US/CorporateCareers"""
import csv, json, re, sys, time, urllib.request, urllib.robotparser

UA = "my-job-export/1.0 (contact: [email protected])"  # say who you are

def parse(url):
    m = re.match(r"https://([a-z0-9-]+)\.(wd\d+)\.myworkdayjobs\.com/(?:[a-z]{2}-[A-Z]{2}/)?([A-Za-z0-9_.-]+)", url)
    if not m:
        sys.exit("Not a myworkdayjobs.com career-site URL")
    return m.groups()  # tenant, data centre, site

tenant, wd, site = parse(sys.argv[1])
host = f"https://{tenant}.{wd}.myworkdayjobs.com"
api = f"{host}/wday/cxs/{tenant}/{site}/jobs"

robots = urllib.robotparser.RobotFileParser(f"{host}/robots.txt")
robots.read()
if not robots.can_fetch(UA, api):
    sys.exit(f"robots.txt of {host} does not allow {api}")

jobs, seen, offset, total = [], set(), 0, None
while True:
    body = json.dumps({"appliedFacets": {}, "limit": 20, "offset": offset, "searchText": ""}).encode()
    req = urllib.request.Request(api, data=body, headers={
        "Content-Type": "application/json", "Accept": "application/json", "User-Agent": UA})
    with urllib.request.urlopen(req, timeout=30) as r:
        page = json.load(r)
    if total is None:
        total = page["total"]  # only the first page reports the total
    posts = page.get("jobPostings") or []
    for p in posts:
        if p.get("externalPath") in seen:
            continue
        seen.add(p.get("externalPath"))
        jobs.append({
            "title": p.get("title"),
            "jobId": (p.get("bulletFields") or [""])[0],
            "location": p.get("locationsText"),
            "posted": p.get("postedOn"),
            "url": host + "/" + site + p.get("externalPath", ""),
        })
    offset += 20
    if not posts or offset >= min(total, 2000):  # the list stops at 2,000
        break
    time.sleep(1)  # one request a second

json.dump(jobs, open("jobs.json", "w"), indent=2)
with open("jobs.csv", "w", newline="") as f:
    w = csv.DictWriter(f, fieldnames=list(jobs[0]))
    w.writeheader()
    w.writerows(jobs)
print(f"{site}: {total} listed, {len(jobs)} exported")
Command
python3 workday_jobs.py https://mastercard.wd1.myworkdayjobs.com/en-US/CorporateCareers
head -3 jobs.csv
Output (the run took 2 minutes 17 seconds)
CorporateCareers: 1037 listed, 1037 exported
title,jobId,location,posted,url
"Manager, Software Engineering",R-279658,"Ramat-Gan, Israel",Posted Today,https://mastercard.wd1.myworkdayjobs.com/CorporateCareers/job/Ramat-Gan-Israel/Manager--Software-Engineering_R-279658
"Manager, Accounting",R-291473,"Pune, India",Posted Today,https://mastercard.wd1.myworkdayjobs.com/CorporateCareers/job/Pune-India/Manager--Accounting_R-291473

The script does not split by facet, so on sites above 2,000 jobs it saves the first 2,000; add the facet loop from above if you need them all.

What Workday's public list does not give you

  • No department or team. The job category facet is the closest thing, and only as a filter.
  • No structured pay. Where an employer publishes a range, it is inside the description text.
  • No company name. You know the tenant (mastercard), not a display name.
  • Coarse dates ("Posted 30+ Days Ago") unless you read each job's detail.
  • Some sites do not answer scripted requests at all. Treat that as a no and move on.

When a hosted tool is easier

One career site, once, is a script. It gets harder when you watch thirty companies every morning and want only what is new: you need a schedule, a place to keep yesterday's ids, retries on HTTP 429 and 5xx, and the same columns for every company. That is what the Workday Jobs Scraper on the Apify Store does with the endpoint above. It returns one uniform row per job, reads up to 2,000 jobs per career site per run (it does not split by facet), and its new-only mode returns only jobs that were not there on the previous run. Run it on Apify's scheduler, or call it from the Apify API.

Tinlark is not affiliated with Workday, Inc., Mastercard, NVIDIA or Salesforce. Their names appear only to point at their public career sites. The postings belong to the employers: check their terms and the laws that apply to you before you republish them.

Questions before you start?

Write to [email protected]. Each product page lists what the tool does, what it does not do, and its exact price.