[Draft — pending publication]
This finding has editorial sign-off but is not cleared for public publication. Numbers are staged for review against the methodology standards; nothing here is final.
Same Job, Five Months Apart: What One Network's Software Postings Did Not Show
Match a network of large-employer software postings employer-by-employer and title-by-title and two facts sit side by side. Hold the employer mix constant and the average description barely moves — the late-winter and midsummer centroids land closer together than two random halves of either window, well below the drift threshold. Leave the mix in and the raw distance is nearly three times that noise floor, because which employers were hiring shifted. So the honest headline is two-part: within employers the software description did not drift; the employer mix did. What else moves is specific and small — a suggestive tilt toward cloud-native tooling, far more ads that mention AI than require it, a mild nudge from onsite toward hybrid — and the whole comparison straddles two greyed ingest months and a half-finished July, so it is exploratory context, not a seasonal verdict.
July 14, 2026 | the alldone.jobs research desk | Snapshot of July 12, 2026
The Record has learned to distrust the story it wanted to find. Its first finding went looking for a summer hiring collapse in restaurants and came back with a modest, heavily bounded shortfall instead; the discipline of that piece was in what it refused to claim. This is the sequel to that habit, on the question the whole observatory was built to ask. If artificial intelligence is rewriting knowledge work, the software job is where the rewrite should show first and clearest. So we matched the software postings in one large-employer network against themselves — the same employers, the same job titles, late winter against midsummer — and read the descriptions for the drift.
The drift is not there — not in the average description, and not within employers. Cohort-matched and employer-balanced, the mean of what February's software postings say and the mean of what July's say sit closer together than either sits to a random half of itself. The pay is flat, the map barely moves, the required skills are the same list in nearly the same order. But the same test shows one thing moving, and honesty requires naming it in the same breath: which employers were posting shifted between the windows, enough that the raw, unbalanced distance runs nearly three times the noise floor. That movement is real — it is composition, a change in the mix of companies hiring, not in the job they describe. So the headline is two-part. Within employers, the software description did not drift beyond noise; the employer mix did move. And a caution rides on the first half: the measure that returns the null is a single averaged vector over some 50,000 templated postings, which can only register a wholesale rewrite at corpus grain — it is blind by construction to change that is sparse, localized, or self-cancelling, the kind the finer skill, AI-tier and term instruments below are built to probe. Over these five months, in this network, the software description did not undergo the loud rewrite. The quieter movements it might hide are exactly what the rest of this piece measures.
What this comparison is, and what it is not
The design is a within-role match, chosen so that a change in who is hiring cannot masquerade as a change in the job. We took every posting coded to the federal computer-and-software occupation family (SOC 15-12xx), normalized each job title by a deterministic rule, and kept only the roles that appear in both windows under the same employer — the same company posting the same normalized title, late winter and midsummer. That intersection is 6,945 matched employer-and-title keys, drawn from 47,764 keys in the earlier window and 33,759 in the later one, and it carries 51,211 postings from Window A and 46,185 from Window B. Every number below is a comparison inside that matched cohort, not across two open populations. Within-employer matching does not remove every composition effect — an employer can post five of a title in February and fifty in July, and those volume shifts still move a pooled share — but it dissolves the largest one, and it is the reason a "same-description" reading is worth trusting at all.
Two facts about the windows have to travel with every sentence that follows, because they constrain the whole comparison rather than qualify a corner of it.
The premise straddles greyed months. This publication's window policy (methodology §13) marks the early corpus as documented context, never load-bearing: February 2026 is not a stable baseline — some 87% of its in-window postings land on a single ingest day — and March is an ingest burn-in backfill, its composition shaped by the pipeline coming online rather than by hiring. The later window's July is a partial month, cut off at the July 12 snapshot. Window A here is exactly those greyed months (February–March); Window B is June with a half-finished July. So this comparison is, by construction, exploratory. It is not a published seasonal claim, and no headline here rests on February being February. What makes it worth reporting is the direction of the result: most of the confounds in these windows push toward spurious difference, not away from it, and we found little — with one exception, the load-bearing semantic measure, whose main confound runs the other way and which is defended on its own terms in the methodology, not by this headwind logic.
The two windows are deduplicated very unevenly. The corpus fingerprint used to collapse repeated postings covers only about 12.9% of February and 11.7% of March postings, but 38.6% of June and 99.9% of July. The consequence is visible in the collapse rate: Window A's raw postings shrink only 3.26% when reduced to distinct content, while Window B's shrink 40.44%. Near-duplicate template text therefore survives in the earlier window and is stripped from the later one. Any comparison of shares or quantiles between the windows — pay, work arrangement, skill counts — may partly reflect that asymmetry rather than real role change, so each such reading below carries the caveat with it, and both raw and distinct counts are reported so the gap stays visible. The pooled cohort clears this publication's suppression rules (§7): no window is dominated by a single employer, whose largest share is 10.7% in Window A and 12.9% in Window B, and individual cells with fewer than thirty postings are flagged and withheld from any factual reading.
No wholesale rewrite — and a composition shift underneath
The load-bearing measurement is the one that is hardest to fake, and it is the one that returns the clearest null. For every posting we take the description, strip the employer boilerplate, and embed the remaining role text as a vector; the centroid of those vectors is a summary of what the window's postings mean. If the software job were being rewritten between the windows, the two centroids would pull apart. The honest test is not whether they differ at all — any two samples differ a little — but whether they differ by more than the window's own internal noise. So we split each window into random halves and measured how far apart those halves land: that within-window distance is the floor, the amount of movement that means nothing. One property of the instrument has to be stated before the result, because it bounds what a null can mean: the centroid is a single average vector over roughly 50,000 heavily templated postings, so it moves only when the whole window's language moves together. It registers a wholesale rewrite; it is structurally insensitive to change that is sparse (a new responsibility in a few percent of ads), localized to one cluster of roles, or offsetting (half the postings drifting toward AI, half toward legacy, leaving the mean fixed). It is the right tool for the loud version of the question and the wrong one for the quiet version — which is why the skill, AI-tier and term readings later carry their own, finer weight.
Held within employers, the cross-window distance sits below that floor; left raw, it sits well above it. Both are in the table, because both are true.
| Measure | Value |
|---|---|
| Cross-window distance, employer-balanced (the design's measure) | 0.000027 |
| Cross-window distance, raw (employer mix left in) | 0.000226 |
| Within-window noise floor (random half-splits, stochastic) | ~0.000078 |
| Balanced ÷ floor | ~0.35 — below the 2.0 gate |
| Raw ÷ floor | ~2.9 — clears the 2.0 gate |
| H1 guardrail on the balanced (design) measure — drift real only if ratio ≥ 2.0 | FAIL |
| Vector coverage, Window A / Window B | 99.8% / 99.7% |
From the R7 cohort exhibit, snapshot of July 12, 2026.
Read the two ratios plainly, because the table is telling two things at once. The design's measure — the one that holds the employer mix constant to ask whether the description itself drifted — puts the February-to-July movement at about 0.35 times the within-window noise, well under the 2.0 gate and in fact below the gap between two random halves of a single window. Leave the employer mix in, and the raw distance runs about 2.9 times that same floor — it would clear the gate. Both numbers are real. Which one answers the question is the design's choice, and it is a definitional one, stated plainly: this piece asks whether the software job's description drifted within employers, so the balanced measure is the answer, and the answer is a null. The raw excess is not error to be waved away — it is a genuine, different phenomenon, and the next paragraph is about it. (The floor itself is stochastic: it is estimated by randomly halving each window, and the exhibit reports the two half-split baselines as 0.000076 and 0.000078 with standard deviations of ±6–8×10⁻⁶, so its last digit and the second digit of each ratio wobble from run to run — an independent re-run lands the balanced ratio nearer 0.36. What does not move is the ordering: on every run the balanced distance sits far below the floor and the raw distance well above it.)
That raw excess has a name, and it is the second finding of this piece, not a nuisance to be scrubbed. The raw cross-window distance — before the employer mix is balanced out — is 0.000226, about eight times the employer-balanced 0.000027. Roughly seven-eighths of the apparent winter-to-summer movement, then, is not the job's description changing; it is the set of companies posting changing — which employers were in the market for software people shifted between the windows. That is composition, and it is real: an economist who counts who is hiring as part of what a job market is would read it as a genuine move. It is simply a different move from the one the headline measure isolates, and the point of the cohort design is to separate the two rather than blur them into a single "distance" — not to make the raw number disappear. A word-level version of the same test agrees in direction, with the caveat that its exact counts do not reproduce: the top-twenty term list is drawn from a non-deterministic deduplication pass, so which terms clear the cut shifts run to run and the precise tallies are not publishable as figures. Read qualitatively, roughly half of the twenty terms that rise most from winter to summer are genuine role content, and the single largest riser is "ai"; fewer than half of the twenty falling terms read as real role language, the rest being template residue and boilerplate — "com," "company," "office," "work". (The face-validity classifier is deliberately conservative and carries no URL flag, so it counts bare web fragments like "https" and "www" as role content rather than boilerplate — which, if anything, inflates the role-content tally on the falling side rather than the reverse.) The terms leaving the corpus are mostly the scaffolding the earlier, less-deduplicated window carried, not skills leaving the job.
A null is not proof of a frozen job. It is the narrower statement that the average description, held within employers, did not move beyond this instrument's within-window noise — in a comparison whose only clean primary-window month is June, set against greyed February and March and a July truncated at the snapshot on the 12th. A change too small, too slow, too localized, or too self-cancelling across postings to move the mean would read here as silence, because a centroid over some 50,000 templated docs can only see a wholesale, corpus-grain rewrite. What the result rules out is exactly that loud version — a wholesale rewriting of the software description in this network over these months — and no more. It does not, on its own, rule out the quiet versions; the skill, AI-tier and term readings are where those are chased.
The pay is flat, and the map barely moves
The softer measures corroborate the null, each with the uneven-dedup caveat attached because each is a share or a quantile.
Compensation is flat through the middle of the distribution. Across the postings that disclose a pay range — about 25,000 in Window A and 16,000 in Window B — the median posted midpoint moves from \$144,350 to \$145,200, a rise of 0.6%, while the mean slips about 1.4% (\$144,704 to \$142,738) — the two disagree because the tails compress: the tenth percentile falls about \$5,800 (\$83,795 to \$78,000) and the ninetieth about \$4,600 (\$206,150 to \$201,600). These are posted range midpoints, not realized pay, drawn from a self-selected subset of employers who choose to disclose, and the mild tail movement is exactly the kind of figure the dedup asymmetry can nudge. Two further cautions sit under these quantiles. They are pooled and not employer-balanced — unlike the semantic test, comp inherits both the dedup asymmetry and the between-employer composition confound the balancing removed, which makes it one of the least confound-hardened measures here, weak corroboration rather than independent confirmation. And the comp-disclosing subset is more concentrated than the cohort at large: its single largest employer holds about 20% of the disclosing postings in each window, against the 10.7%/12.9% of the full cohort — still under this publication's §7 majority bar, but stated so "flat pay" is not read as resting on the same broad base as the rest. With those attached, the finding is the flat median, not the tails: software pay in this cohort did not move over the window.
Geography barely moves — with one artifact to name first. The single largest state-level cell movement is not a state at all but the unparsed "(none)" bucket, postings whose location did not resolve, which rises 1.5 points (28.7% to 30.2%) — a parsing-and-coverage artifact, larger than any real geographic move, and set aside. Among named states the largest shift in either direction is under two-thirds of a percentage point: Texas rises 0.6 point (4.2% to 4.8%) and Washington falls 0.6 point (3.1% to 2.5%), while California (5.8% to 6.1%), Virginia (5.1% to 5.3%) and New York (2.6% to 2.6%) are effectively unchanged. The metro picture is the same near-null. One metro, Raleigh–Cary, drops from 388 postings to zero across the windows — but that is a feed going quiet, not a hiring event, and it is precisely why no single geographic cell is allowed to carry a claim here. Employer size mix is likewise stable: the largest employers hold 89.0% of the cohort in winter and 88.1% in summer, a 0.9-point drift. Same pay, same map — and, holding the employer mix constant, the same description.
What did move, precisely
Three things move enough to name, and each gets modest, bounded treatment.
A suggestive tilt toward cloud-native tooling. Read the required-skills list and the same technologies rise together in a pattern that reads like real practice: python up about three and a half points on both the mentioned-skill and required-skill measures (+3.5 and +3.7 points), with sql, java, go, kubernetes, terraform, linux, rust and docker all rising, and javascript, c/c++ and gcp easing back. That is a recognizable direction — toward backend, infrastructure and systems tooling, and away from front-end — but it should be read as a direction only, and a suggestive one. Nearly every technical skill rises on the summer side at once: python, sql, java, go, kubernetes, linux, terraform, rust, docker, aws, troubleshooting and documentation all move up together. A near-universal technical rise is the fingerprint of the skill tagger simply catching more on the fresher, far-more-deduplicated July postings — tagging-coverage inflation, not necessarily rising demand. What survives that caveat is the pattern — cloud-native and systems tooling gaining relative to front-end — and not the magnitude of any single mover.
It carries one firm caveat, disclosed rather than buried. The single largest apparent skill mover is not a real one: "communication" rises 8.7 points while "communication skills" falls 4.3 points and "strong communication" falls too — the same idea splitting and merging as the skill tagger's canonical form drifts between windows. The same thing happens to the continuous-integration tooling, where "ci cd pipelines" (+3.6), "cicd pipelines" (−2.4) and "ci/cd" (+1.8) are one concept churning across three spellings. Those variant pairs are tokenization artifacts, not demand changes, and we set them aside. The coherent, non-variant movers survive that variant filter, but they still sit under the coverage caveat above — which is why the tilt is offered as a suggestive direction, its pattern (cloud-native over front-end) the part we trust, its magnitude not.
More ads talk about AI than require it — and that gap is the finding. Split every posting into three tiers of AI language and the tiers move in different directions. The share that merely mentions AI climbs 5.6 points, from 54.5% to 60.1% — by summer, three in five software ads in this cohort name AI somewhere in the text. But the share that requires a concrete AI skill barely moves, up 0.9 point to 26.7%, and the share where AI is the core of the role actually falls 0.8 point, to 2.6% — roughly one posting in forty.
| AI language tier | Window A | Window B | Change |
|---|---|---|---|
| Mentions AI at all | 54.5% | 60.1% | +5.6 pt |
| Requires an AI skill | 25.8% | 26.7% | +0.9 pt |
| AI is the core of the role | 3.4% | 2.6% | −0.8 pt |
From the R7 cohort exhibit, snapshot of July 12, 2026.
The three tiers are the independent regex flags the corpus computes for every posting (enrichment/ai_flags.py): mentions AI is any AI or generative-AI reference at all, employer boilerplate ("we are an AI-powered company") included; requires an AI skill is a concrete AI competency or named tool asked for as a requirement — an ML or LLM technique, langchain/RAG/a vector database, PyTorch/TensorFlow, a named model, or an "experience with generative AI" clause; AI is the core of the role is title-driven, reserved for jobs that are AI jobs (an AI/ML engineer, a prompt engineer, an applied scientist with ML context). They are three independent booleans, not a strict nesting, so a posting can trip the broad tier without the narrow ones.
The loudest tier is the one this publication's own method flags as boilerplate-exposed — "the text mentions AI," not "the role needs it" — and it is the only tier that moves. Being a share, that +5.6-point move also carries the same uneven-dedup caveat as every other share here — stale non-AI boilerplate retained in the earlier window depresses its mention count while fresh July uniques lift it — so read the gap between the tiers, not the precise 60.1%, as the finding. This is the same shape The Record found across job ads generally: the growth is in AI being named, not in AI being required. On the software job specifically, over these five months, the descriptions are absorbing AI mostly at the surface.
A mild nudge from onsite toward hybrid. Onsite work arrangements fall 4.4 points (77.1% to 72.7%) and hybrid rises 5.3 points (8.2% to 13.4%), with fully remote nearly flat (+0.8 point). It is a small, single-direction shift, and it is one of the readings most exposed to the dedup asymmetry — a "flexible" category collapsing from 2.3% to 0.7% in the same table looks more like tagging than like policy — so it is reported as a mild tilt, not a return-to-office reversal.
The column we threw out, and why we are showing you
One measure looked like a dramatic finding and is not one, and the honest move is to show you the discard rather than hide it. Sorted by seniority, the cohort appears to lurch from mostly mid-level to mostly senior: "mid" falls 53 points (72.6% to 19.6%) while "senior" rises 20.9 points and postings with no seniority tag rise 19 points. Taken at face value that would be the largest movement in the entire analysis.
It is an artifact, and a knowable one. The seniority derivation changed across the burn-in boundary: "mid" accounts for 62.6% and 68.0% of the two earlier months but only 19.9% and 6.5% of the two later ones, while the untagged share climbs from about 2–3% to 17–23%. Within a matched cohort — the same employer, the same title, present in both windows — the same posting cannot genuinely change seniority class this much; a classifier that relabels two-thirds of "mid" postings between winter and summer is measuring itself, not the labor market. The measure's own guardrail marks it FAIL. So seniority is quarantined: it appears in this article only as a disclosed, discarded column, and no sentence here reads it as a signal. Naming the discard is the point — the same bar that passes the coherent skill tilt is the one that rejects this.
Methodology
Every number above is drawn from a single corpus snapshot dated July 12, 2026. The deterministic aggregates — the cohort counts, the shares and quantiles, and the employer-balanced centroid distances — reproduce from it exactly. Two quantities are stochastic and reproduce in direction rather than to the last digit, and both are flagged where they appear: the within-window noise floor (and therefore the cross-over-within ratio), which is estimated by a random half-split of each window, and the word-level term-movement counts, whose top-twenty term list is drawn from a non-deterministic deduplication pass. The exhibit is analysis/semantic-drift/output/r7/r7_software_cohort.md and its companion .json, and the provenance tag on each figure in the source of this article names the block it comes from. A later revisit would ship as a new article at its own vintage, under this publication's vintage rule, not as a silent edit here.
The design is a cohort match: postings coded to SOC 15-12xx, with each title normalized by a deterministic rule (lowercased; parentheticals and brackets dropped; non-alphanumeric characters collapsed to spaces; seniority tokens kept), and only employer-and-title keys present in both windows retained. The semantic test embeds boilerplate-stripped role text with a small offline sentence model, balances the centroids across employers so a shift in the posting mix cannot register as drift, computes distance only on the distinct-content set, and interprets that distance only against the within-window half-split floor — never as an absolute threshold, because the model's distances are small in absolute terms by construction. The full standing method — flow rather than stock, share of total rather than raw counts, the primary-window and greyed-context policy, suppression below thirty postings and above a single-employer majority, and the single-network scope limit — is set out on the methodology page, which this article is written to be checked against rather than taken on faith.
Two properties of this particular comparison deserve naming here. First, the cohort design is visibly doing its job: the raw winter-to-summer description distance is about eight times the employer-balanced one, so most of the apparent movement is between-employer composition, which the balanced measure — the one we report as the description null — sets aside (and which the article reports separately as its own finding). Second, the two remaining confounds cut in different directions for different measures, and we do not claim a single "headwind" across all of them. For the share and quantile readings — pay, arrangement, skill and AI counts — the earlier window's retained near-duplicate templates and the burn-in classifier changes both tend to manufacture apparent difference, so a null there is reached against a headwind. But for the load-bearing semantic distance the sign plausibly reverses: the near-duplicate role-text that survives in the less-deduplicated February–March window drags that window's centroid toward its own shared template mass, which can compress the cross-window distance rather than widen it. So the semantic null is not defended as "a null against a headwind"; it rests on the employer-balancing and on sitting well below the within-window floor, and we state its direction honestly rather than claim the confound helped us.
Limitations
- The premise straddles greyed context months, by construction. Window A is February–March 2026, which this publication treats as documented context and never as a baseline: February concentrates ~87% of its postings on one ingest day and March is an ingest burn-in backfill. Window B's July is a partial month, snapshot-truncated at July 12. This comparison is therefore exploratory, not a published seasonal claim; its value rests on the direction of the result and on the confound-hardened design — for the share-based readings, a null reached despite windows that tend to manufacture difference; for the load-bearing semantic measure, a null defended on its own terms in the methodology (its main confound plausibly runs the other way) — not on the calendar identity of either window.
- The two windows are deduplicated very unevenly. The content fingerprint covers ~12.9% of February and ~11.7% of March but ~38.6% of June and ~99.9% of July, so distinct-content collapse is 3.26% in Window A against 40.44% in Window B. Every share or quantile comparison — pay, work arrangement, skill counts, AI tiers — may partly reflect that asymmetry rather than real role change. The load-bearing semantic test is run on the distinct-content set with employer-balanced centroids to blunt this, but the softer readings carry it in full.
- The null does not prove a frozen job. The semantic result rules out only movement of the average description, held within employers, larger than this instrument's within-window noise floor — and it does so in a comparison whose sole clean primary-window month is June (February and March are greyed context; July is truncated at the 12th). A change too small, too slow, too localized, or too self-cancelling across postings to move the mean would read here as silence, because a centroid over ~50,000 templated docs registers only a wholesale, corpus-grain rewrite, not sparse or localized change. The claim is bounded to exactly that: no wholesale rewrite of the software description in this network over these months; the finer skill/AI-tier/term instruments are the probes for the smaller tilts.
- Scope is one applicant-tracking network. Every claim is scoped to direct employer postings from large U.S. employers (the DirectEmployers network), never "the labor market" or "the software workforce." This is a within-corpus comparison of that record against itself.
- Skill movers carry canonicalization and coverage caveats. The largest apparent skill shifts are tokenization variants of one concept ("communication" vs "communication skills"; "ci cd pipelines" vs "cicd pipelines" vs "ci/cd") and are set aside. The coherent, non-variant tilt (python, sql, java, go, kubernetes, terraform, rust, docker up; javascript, c/c++, gcp down) is reported as a direction, not a magnitude, because a skill's share can rise from better tagging coverage as well as from real demand.
- Compensation is a disclosed-subset midpoint. Comp quantiles are computed over the ~25,000 (Window A) and ~16,000 (Window B) postings that disclose a range, are posted range midpoints rather than realized pay, and are sensitive to which employers choose to disclose. The reported finding is the flat median (+0.6%); the mild tail compression is within the range the dedup asymmetry can move.
- Seniority is a confounded, quarantined column, not a finding. The seniority derivation changed across the burn-in boundary — "mid" falls from ~63–68% of the earlier window to ~7–20% of the later one while untagged postings climb from ~2–3% to ~17–23% — which no matched cohort could genuinely produce. Its guardrail is FAIL; it appears only as a disclosed discard and carries no signal.
- Matching grain is presence, not paired reweighting; balancing is at employer grain. A key is retained if the employer-and-title pair appears in both windows. The employer-balancing that isolates the semantic null operates at employer granularity — it equalizes which companies contribute, and so removes only the between-employer slice of composition. It does not touch shifts within a single employer: an employer posting five of a title in February and fifty in July still volume-weights the pooled shares, and an employer reallocating across its own titles (its title mix) is unbalanced. So the design dissolves the largest, between-employer confound, not every one; a stricter per-key equal-weight design is future work.
- Not cleared for publication. These numbers are staged for the three-lens adversarial panel (independent reproduction, hostile methodology review, data integrity) and Fable sign-off, the same gate every finding on this site passes before it runs. Nothing here is final.