Codex, working the defensive half without having written the parser, probed M1 with adversarial input and found defects my own tests missed. Verified independently before fixing, and two more were found while confirming: negative_array_count ACCEPTED as an empty array negative_hash_count ACCEPTED as an empty hash negative_bignum_words NoMethodError leaked outside StreamError negative_ivar_count ACCEPTED negative_string_len cursor moved BACKWARD, wrong error raised The last is the worst. take(-5) does not trip the count > remaining guard, so byteslice returns nil and @position decreases. A parser whose cursor can rewind on attacker input is a loop primitive, not merely a wrong error. The M1 gate claimed bignum length confusion was covered. It was not. Oversized widths were tested and negative counts never were, because the same author chose both the implementation and the cases it would face. That is the negative-control failure one level up, and it is exactly what an author cannot catch alone. Fixes: every count and length now flows through read_count with a role label and a nonnegative check, take rejects negative byte counts outright, and MalformedCountError joins the StreamError hierarchy so nothing leaks a raw NoMethodError. Negative link indices were already guarded; regression tests now pin that. Also closes a detection blind spot Codex identified. read_ivar and read_object discarded the parsed name nodes after taking their values, so a sink tag placed in an instance-variable-name position vanished from Result#sinks. Node now carries an auxiliary collection that Node#each traverses, and name and class nodes are retained. Proven: a stream with a userdef tag in the name position now reports Evil#_load where it previously reported nothing. 46 parser tests, 80 assertions. All controls pass, exploit gate still passes. |
||
|---|---|---|
| .. | ||
| lib | ||
| scripts | ||
| target | ||
| test | ||
| .gitignore | ||
| .rubocop.yml | ||
| CHANGELOG.md | ||
| Gemfile | ||
| LICENSE | ||
| README.md | ||
| Rakefile | ||
| justfile | ||
| rube.gemspec | ||
README.md
rube
A Ruby object-deserialization security lab.
A gadget chain is a Rube Goldberg machine. One untrusted blob goes in, a dozen unrelated standard-library methods knock each other over, and code execution falls out the far end. This project builds the machine, then builds the thing that stops it.
Why this exists
Marshal.load on untrusted input is arbitrary code execution. So is YAML.unsafe_load,
JSON.load with additions enabled, and Oj.load in its default mode. This is not a Ruby
quirk. It is the same class of bug as Java deserialization, PHP POP chains, and Python
pickle, and it sits at CWE-502 in the CISA Known Exploited Vulnerabilities catalog with a
34.8% known-ransomware rate against a 20.1% baseline across the catalog as a whole.
Most write-ups on this topic teach the exploit. Fewer teach why the obvious defense does not work. This one does both, because the second half is where the actual lesson lives:
You cannot make Marshal.load safe with an allowlist. The proc you pass runs in
r_post_proc, which marshal.c invokes after load_funcall(... s_mload ...). By the
time your allowlist sees the object, marshal_load has already run. The pattern widely
copied off Stack Overflow is a post-mortem, not a veto.
Psych's allowlist genuinely is a veto — for exactly one reason. It checks the tag before revival, where Marshal checks the object after construction. Identical intent, opposite outcome, decided entirely by where the check sits.
Status
Under construction. What exists and is tested:
- Marshal stream parser — parses the binary format, extracts referenced class names
and gadget sinks, and validates structure, all without ever calling
Marshal.load. Rejects truncated streams, unsupported versions, unknown tags, out-of-bounds object links and symlinks, oversized fixnum widths, trailing bytes, and excessive nesting.
Planned: version-compatibility matrix, reflection-based gadget scanner, payload builder, a deliberately vulnerable containerized target, and the defensive layer.
Usage
require "rube"
payload = Marshal.dump(Gem::Requirement.new(">= 0"))
result = Rube::Marshal::Parser.new(payload).parse
result.class_names
# => ["Gem::Requirement", "Gem::Version"]
result.sinks.map { |s| "#{s.class_name}##{s.sink_method}" }
# => ["Gem::Requirement#marshal_load", "Gem::Version#marshal_load"]
Nothing above instantiates a class, calls a constructor, or invokes Marshal.load.
Development
Everything runs in Docker against a pinned Ruby.
just test run the parser suite
just control run the negative controls
just check both
just build build the gem with --strict
just manifest list exactly what would ship in the .gem
A note on the object-link index
Ruby's Marshal format documentation states that object links are one-indexed. They are
zero-indexed. A self-referential array dumps as 04 08 5b 06 40 00, where the trailing
00 is a link to the outermost object at index 0. The parser is written against the
observed bytes, not the documentation.
License
AGPL-3.0-or-later. See LICENSE.