Skip to content

[pull] main from daijro:main - #24

Merged
pull[bot] merged 12 commits into
scraperapi:mainfrom
daijro:main
Sep 28, 2026
Merged

pull[bot] merged 12 commits into
scraperapi:mainfrom
daijro:main

Conversation

@pull

@pull pull Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

See Commits and Changes for more details.


Created by pull[bot] (v2.0.0-alpha.4)

Can you help keep this open source service alive? 💖 Please sponsor : )

coffeegrind123 and others added 12 commits September 27, 2026 20:09
pythonlib left the model to fpgen, whose first import downloads the
first release the GitHub API lists: model-4/2025, whose WebGL records
have no vendor or renderer. Every generated launch on a fresh install
failed with KeyError: 'vendor'. scripts/pin-fpgen-model.py pinned the
model for CI only, and fpgen's five-week refresh replaced even that
pin, under a stamp the script's --check still trusted.

fpgen is now imported only through fpgen_model.load_fpgen(), which
installs the release named by the pin, checks the archive and each
file against their sha256, and dates the files past fpgen's refresh.
It writes the layout and stamp the script and the TypeScript launcher
use, verifies an install by hashing it, and leaves FPGEN_MODEL_URL to
fpgen. `camoufox fetch` installs the model as well. The pin gains each
file's sha256; the package carries a copy of it, and the script now
calls the module.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
(cherry picked from commit f043de3)
Shipping distribution/policies.json puts Firefox in enterprise-policy
mode, and since Firefox 151 that mode sets the default of
dom.webserial.enabled to false (EnterprisePoliciesParent.sys.mjs). So
navigator.serial, Serial and SerialPort are missing, while stock
Firefox 152 exposes them on secure pages (the pref defaults to true on
desktop). A page can see the difference.

DefaultSerialGuardSetting: 3 sets the default back to true without
locking it (Policies.sys.mjs). stock-parity-probes.py now checks that
navigator.serial exists on its secure-context probe page.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 6f5e1e4)
browser-init.patch installs the `addons` entries from gBrowserInit,
which runs for every browser window, and Juggler opens a window for
every page of every new context. Reinstalling a temporary addon
restarts it: uBlock Origin restarted with nearly every new context,
and a navigation its webRequest listener had suspended when a restart
hit was never resumed, so page.goto timed out. Addons now install once
per launch, flagged by a default-branch pref, which is never written
to prefs.js.

tests/patches/addons-install-once.py opens 12 contexts, one page each,
and counts uBO background-page loads from the DocumentChannel log.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
(cherry picked from commit c23207c)
Juggler set inRDMPane on every page with an emulated viewport, which is
Playwright's default for new contexts. RDM is devtools' mobile mode, and
a page can read it: scrollbars become overlay scrollbars with no layout
width, so with classic scrollbars pinned a page measured 12 px without a
viewport and 0 px with one. Navigator, screen and window getters also
take RDM branches. The viewport itself is sized by the browser element
and does not need RDM.

Playwright's Juggler has enabled RDM only for isMobile since
microsoft/playwright#41859. Do the same: Browser.setDefaultViewport and
Page.setViewportSize already accept isMobile, and it is now kept and
applied instead of dropped, so is_mobile=True still gets RDM.

The new guard tests/patches/viewport-no-rdm.py fails on the old Juggler
(0 px in a viewport context against 12 px without) and passes on this
one. It also checks that is_mobile=True still turns RDM on.

Five upstream Playwright tests go in ci/skiplist.yml. They assume
headless scrollbars take no width, which upstream gets by hiding them
with a style sheet that Camoufox removed. They only passed here because
of RDM's overlay scrollbars.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 77d69ac)
coherence rejected "Intel(R) HD Graphics 400" and "Radeon R9 200
Series" on macOS as a Braswell Atom IGP and a desktop PC card. They are
Firefox's sanitized buckets, not devices: SanitizeRenderer
(FIREFOX_152_0_4_RELEASE) maps an Intel Mac's UHD 630 to the first and
a Radeon Pro 5300M to the second, per Firefox's own TestCiMac and
TestMacAmd. fpgen's pinned model records both from Firefox on macOS at
1.14% each. The rule kept every generated Mac off them, removed 24
real macOS presets in #779 (restored here, 9 + 15), and dropped either
GPU from a caller's Mac preset. ANGLE and llvmpipe stay rejected:
Firefox 152 has no ANGLE-on-Metal path (Bug 2046027 came later).

Intel Macs are now checked as machines instead (intel-mac-hardware),
for every non-Apple GPU on macOS:

- a core count some Intel Mac with that GPU reports. Firefox reports
  physical cores where kern.tcsm_available is set and logical ones
  otherwise, so either counts: 2-8 physical / 4-16 logical for the
  IGP, up to the 2019 Mac Pro for a discrete GPU;
- a screen that is not a notched MacBook's or the 24" iMac's. 45.5% of
  fpgen's Firefox macOS screens are one, and the draw already put an
  Intel or AMD bucket behind 4.9% of them.

The WebGL draw applies the same check, given the core count. The preset
GPU drop moves next to the WebGL draw, after the host core count and
the display clamp: before, a preset's GPU was judged against cores the
launch then replaced. The TypeScript launcher gets the same changes,
and the goldens cover the new check and the narrowed draws.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
(cherry picked from commit d6e2adc)
RoverfoxStorageManager::ReadValue fell through its local cache to
Preferences::GetCString on whatever thread asked, and WorkerNavigator
asks from worker threads. libpref's table is main-thread only and a
release build looks it up without a lock, so a worker reading navigator
while the main thread inserted another context's prefs crashed the
content process.

Off the main thread the local cache is now the whole answer. The main
thread keeps it a mirror of every roverfox.s.* pref: a snapshot the
first time the process uses the storage there, then the prefix
callback that already cleared the miss cache.

tests/patches/worker-config-reads.py creates contexts while four
workers read navigator. On v152.0.4-beta.30 it crashed the content
process in 21 of 21 runs; without workers the same loop held for 60 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 338fe9f)
…om the pin

CAMOUFOX_FPGEN_DATA may point the TS launcher at pythonlib's fpgen data/,
and it decompresses values.dat there. ensure_fpgen_model() refused to run
whenever values.dat existed, so sharing the directory broke every Python
launch. The pin now carries values.dat's sha256 (all three twins): a
matching values.dat is kept, any other is removed, since fpgen reads it in
preference to the verified archive. The files `fpgen decompress` leaves
still fail loudly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
When a content process crashes, Juggler disposes the page's target,
which drops it from its browser context. Nothing closed the tab after
that: BrowserContext.destroy() only closes the pages it still tracks,
and the default context closes none. Every crashed page kept its
browser window until the browser exited, 15-60 MB of parent RSS each,
growing without bound across crashes. Playwright's own Juggler has the
same code and the same leak.

The crashed tab is now closed on the next tick after the crash is
handled. A persistent launch has no -silent survival area, so closing
its last window would quit the browser; there the crashed tab is kept
until another tab opens.

test_the_parent_stays_flat_across_content_crashes passed while leaking.
#785's warm-up moved its baseline past the first context's 150-200 MB,
which left three crashes' worth of this leak under the 400 MB bound.
It now compares 8 context cycles with a content crash each against 8
without, after two warm-up cycles, and fails if the crashes add 100 MB
or more. Both numbers print on every run.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
(cherry picked from commit bfd9398)
Responsive Design Mode plus has_touch swallowed a mouse click's pointer
events: page.click() fired mousedown, mouseup and click with no
pointerdown or pointerup, which stock Firefox never does. Found while
verifying #798, which fixes it by enabling RDM only for isMobile.
viewport-no-rdm now compares a click in a has_touch context with one on
the launch-level page.

On beta.31 with main's Juggler the guard fails on both checks; with
this branch's Juggler it passes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…obile

#798 left RDM on for isMobile, as Playwright's Juggler does. RDM matches
no real browser: Firefox for Android never runs it, and under touch
emulation it drops a mouse click's pointer events, which no device does.
Camoufox only has desktop identities, so an is_mobile context was a
desktop UA, platform, fonts and GPU with devtools' mobile mode on top.

Juggler now keeps inRDMPane off for every page, is_mobile included, and
the isMobile plumbing #798 added is gone again. viewport-no-rdm requires
is_mobile=True to keep the platform's scrollbars too.

Both launchers warn instead (warnings.yml is_mobile): on new_page() /
new_context(is_mobile=True) of a Camoufox browser, and on is_mobile
passed to launch_options() for a persistent context. has_touch,
device_scale_factor and viewport keep working without RDM.

Upstream playwright-python skips its isMobile tests on Firefox in
1.61-1.63, so no skiplist entries are needed.

On beta.31 with this Juggler in omni.ja, viewport-no-rdm passes: 12/12 px
of scrollbar with no viewport, with a viewport, and with is_mobile, and a
has_touch click fires pointerdown and pointerup. pythonlib 411 passed;
typescript 584 passed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nd oscpu too

fpgen's Linux pool now and then pairs a Linux user agent with
navigator.platform Win32 and a Windows oscpu. launch_options() corrects
that with fix_navigator_arch(); generate_context_fingerprint() never
called it, so 24 of 1,500 Linux NewContext identities (1.6%) said Win32
under a Linux UA -- and because that path reads the OS for fonts and
voices from the platform, they drew Windows fonts and voices as well.

It now applies the same fix right after the fpgen draw, in pythonlib and
in the TypeScript twin. The fix draws nothing, so the parity goldens are
unchanged.

The new tests feed the context path a real Linux draw with the Windows
platform and oscpu, and fail without the fix ('Win32' != 'Linux x86_64')
in both ports. After it: 0 of 1,500 sampled contexts mismatch.
pythonlib 412 passed; typescript 585 passed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… growth on every PR (#789)

* ci: split the patch guards by kind, and run memory growth on every PR

Patch guards: one 11-minute job ran all 26 guards and the Playwright
skiplist audit. Whether the patches apply is the build's check; the
guards test that what they do still works, and fall into three kinds,
now three jobs beside the skiplist audit:

  spoofing    a spoofed value still reaches the page and holds together
  automation  Playwright stays invisible to the page and never deadlocks it
  parity      what a page, or the OS, can observe matches stock Firefox

Each writes its own suite (patch_guards_<group>), so a failure names the
kind that broke. GROUPS in ci/run_patch_guards.py assigns every guard to
exactly one, and a self-test fails on a guard in none. One job id with a
matrix, so everything that needs patch-guards is unchanged.

Memory growth: ~38 minutes in one process kept it on the schedule and
out of the gate. ci.run_native --shard i/n runs every n-th collected
test, and the growth job is a 7-way matrix -- one test per runner, about
six minutes each -- on every pull request, required by the summary and
the gate. summarize.py already folds <suite>-<i>of<n> results back into
one suite, as it does for Playwright.

CONTRIBUTING.md now says why the stealth check skips on a fork pull
request: GitHub gives secrets only to branches in this repository.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci(release): publish to npm after PyPI, from the same commit

The two launchers ship at one version, but were released by two
unrelated, hand-started workflows, so npm could get a release PyPI did
not. "Publish to pypi" is now the one place a release starts:

1. it calls publish-npm.yml as a dry run -- every check, the build, the
   pack check and `npm publish --dry-run` -- so a broken npm package stops
   the release before anything is uploaded;
2. it uploads to PyPI;
3. its success triggers publish-npm.yml (workflow_run), which publishes
   the commit PyPI was released from.

publish-npm.yml stays the file that publishes, because npm's trusted
publisher is tied to its name. Started by hand it only retries the npm
half, and refuses unless PyPI already has the version.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(guards): contentaccessible-parity reads its probe's marked result line

The live probe runs in a child process and the parent parsed its whole
stdout as JSON. On a machine whose cache has no addons yet, the first
Camoufox launch downloads uBlock Origin and prints its progress to
stdout first, so the parse failed ("Expecting value: line 2 column 1").
It only ever passed because another guard launched Camoufox earlier in
the same job; split into its own leg, it ran first on a fresh runner.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci(guards): group the three guards main added since the split

addons-install-once, viewport-no-rdm and worker-config-reads landed on main
after the groups were drawn; the one-group-each self-test caught them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(guards): gfx-probes gives a blocklisted launch a second try

The blocklist signature means gfxInfo is empty. Missing probes cause that on
every launch; a present glxtest that fails or times out on a loaded runner
causes it once in a while, which failed stock parity on this PR. Relaunch once
before failing, and print what the probes wrote to stderr.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
@pull pull Bot locked and limited conversation to collaborators Sep 28, 2026
@pull pull Bot added the ⤵️ pull label Sep 28, 2026
@pull
pull Bot merged commit e92caec into scraperapi:main Sep 28, 2026
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants