Cybersecurity-Projects/PROJECTS/beginner/deserialization-gadget-lab
CarterPerez-dev f5196252e9 feat(rube): M1 Marshal stream parser - inspect payloads without deserializing
Scaffolds the Ruby deserialization security lab and lands its defensive core
first: a parser that extracts structure, referenced class names, and gadget
sinks from a Marshal stream without ever calling Marshal.load.

Sinks are classified along the gated/ungated dispatch axis. Marshal checks
respond_to? before invoking marshal_load and _load, while hash, eql?, <=> and
[]= are dispatched blind, so the same class can be dead as a Marshal entry
point and live as a #hash entry point.

Object links are ZERO-indexed. Ruby's Marshal format documentation says
one-indexed and is wrong: a self-referential array dumps as 04 08 5b 06 40 00
with the trailing 00 linking to the outermost object. Written against observed
bytes rather than the docs.

Validation rejects truncated streams, unsupported version bytes, unknown type
tags, out-of-bounds object links and symlinks, oversized fixnum widths,
trailing bytes, and nesting past a configurable depth limit.

A negative-control script accompanies the suite and caught a test that was
passing vacuously: the TracePoint oracle watched :c_call, but Marshal.load is
a Ruby-level method in Ruby 4.0 (<internal:marshal>:33) and fires :call, so
the test could never have failed. The suite now asserts the oracle observes a
real Marshal.load before the negative assertion is allowed to mean anything.

Gem manifest is an explicit allowlist rather than git ls-files, so the
deliberately vulnerable target cannot be swept into a published gem later.

34 tests, 62 assertions, 0 failures. 52/52 corpus round-trip. gem build
--strict clean. All execution in ruby:4.0-slim with --network none.
2026-07-26 09:32:14 -04:00
..
lib feat(rube): M1 Marshal stream parser - inspect payloads without deserializing 2026-07-26 09:32:14 -04:00
test feat(rube): M1 Marshal stream parser - inspect payloads without deserializing 2026-07-26 09:32:14 -04:00
.gitignore feat(rube): M1 Marshal stream parser - inspect payloads without deserializing 2026-07-26 09:32:14 -04:00
.rubocop.yml feat(rube): M1 Marshal stream parser - inspect payloads without deserializing 2026-07-26 09:32:14 -04:00
CHANGELOG.md feat(rube): M1 Marshal stream parser - inspect payloads without deserializing 2026-07-26 09:32:14 -04:00
Gemfile feat(rube): M1 Marshal stream parser - inspect payloads without deserializing 2026-07-26 09:32:14 -04:00
LICENSE feat(rube): M1 Marshal stream parser - inspect payloads without deserializing 2026-07-26 09:32:14 -04:00
README.md feat(rube): M1 Marshal stream parser - inspect payloads without deserializing 2026-07-26 09:32:14 -04:00
Rakefile feat(rube): M1 Marshal stream parser - inspect payloads without deserializing 2026-07-26 09:32:14 -04:00
justfile feat(rube): M1 Marshal stream parser - inspect payloads without deserializing 2026-07-26 09:32:14 -04:00
rube.gemspec feat(rube): M1 Marshal stream parser - inspect payloads without deserializing 2026-07-26 09:32:14 -04:00

README.md

rube

A Ruby object-deserialization security lab.

A gadget chain is a Rube Goldberg machine. One untrusted blob goes in, a dozen unrelated standard-library methods knock each other over, and code execution falls out the far end. This project builds the machine, then builds the thing that stops it.

Why this exists

Marshal.load on untrusted input is arbitrary code execution. So is YAML.unsafe_load, JSON.load with additions enabled, and Oj.load in its default mode. This is not a Ruby quirk. It is the same class of bug as Java deserialization, PHP POP chains, and Python pickle, and it sits at CWE-502 in the CISA Known Exploited Vulnerabilities catalog with a 34.8% known-ransomware rate against a 20.1% baseline across the catalog as a whole.

Most write-ups on this topic teach the exploit. Fewer teach why the obvious defense does not work. This one does both, because the second half is where the actual lesson lives:

You cannot make Marshal.load safe with an allowlist. The proc you pass runs in r_post_proc, which marshal.c invokes after load_funcall(... s_mload ...). By the time your allowlist sees the object, marshal_load has already run. The pattern widely copied off Stack Overflow is a post-mortem, not a veto.

Psych's allowlist genuinely is a veto — for exactly one reason. It checks the tag before revival, where Marshal checks the object after construction. Identical intent, opposite outcome, decided entirely by where the check sits.

Status

Under construction. What exists and is tested:

  • Marshal stream parser — parses the binary format, extracts referenced class names and gadget sinks, and validates structure, all without ever calling Marshal.load. Rejects truncated streams, unsupported versions, unknown tags, out-of-bounds object links and symlinks, oversized fixnum widths, trailing bytes, and excessive nesting.

Planned: version-compatibility matrix, reflection-based gadget scanner, payload builder, a deliberately vulnerable containerized target, and the defensive layer.

Usage

require "rube"

payload = Marshal.dump(Gem::Requirement.new(">= 0"))
result = Rube::Marshal::Parser.new(payload).parse

result.class_names
# => ["Gem::Requirement", "Gem::Version"]

result.sinks.map { |s| "#{s.class_name}##{s.sink_method}" }
# => ["Gem::Requirement#marshal_load", "Gem::Version#marshal_load"]

Nothing above instantiates a class, calls a constructor, or invokes Marshal.load.

Development

Everything runs in Docker against a pinned Ruby.

just test       run the parser suite
just control    run the negative controls
just check      both
just build      build the gem with --strict
just manifest   list exactly what would ship in the .gem

Ruby's Marshal format documentation states that object links are one-indexed. They are zero-indexed. A self-referential array dumps as 04 08 5b 06 40 00, where the trailing 00 is a link to the outermost object at index 0. The parser is written against the observed bytes, not the documentation.

License

AGPL-3.0-or-later. See LICENSE.