Skip to content

Improvements over bce4 #2

Description

@akamiru
  • Completly vectorized: 16 "butterflies" each iteration
  • Added fast ANS encoder and decoder
    • Reduced compression compared to the adaptive
    • A LOT faster able to keep up with the new level of performance
  • Changed iteration layout
    • Handles 8 queues interleaved instead of one after another
      • Greatly improved cache locality: -75% cache misses
    • Allows caching of rank queries and part of the queues
      • Greatly reduces the number of gathers and scatters: -50% µops
      • Reduces memory usage of the queues by 57%
  • Improved rank update
    • Clears and sets bits base on butterfly variables rather then lzcnt/tzcnt moving bits
    • Updates the prefix sum (base rank) only iff necessary rathe rthan recalculating it every time
    • Never needs to gather the prefix sum (base rank)
    • Logarithmic conflict resultion reduces writes to memory

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions