Passing around an IP address in a single register. Today's CPUs are powerful enough to not need hardware forwarding chips to route up to 1Gbps of traffic easily, without much custom software optimisation. However that only holds true for IPv4 traffic. Throw in IPv6 traffic and suddenly each compare takes twice as much time in the CPU. Functions cannot pass addresses and masks in the registers, requiring stack juggling or worse, memory access.
> Throw in IPv6 traffic and suddenly each compare takes twice as much time in the CPU
So an operation responsible for less than 0.1% of the total time spent routing a packet now takes twice as much time... or it would if CPUs weren't superscalar for a long time now. That's paid back in tenfold just by the omission of the header checksum.
Depends on the ISA and optimization level - and the OS. Win32's x86-only __stdcall convention always passes arguments on the stack instead of in registers.
Ask yourself who benefits from this...slower networking on commodity hardware. Then look at who had the most influence over ipv6 in the IETF, etc. No big surprises.