It inspired me to ask whether there's merit in a purely data driven approach to find and count the unique preference strings for each state - these can then be compared to the HTVs but will include non-HTV patterns such as donkey voting.
==> NT_table.tab <==
Preferences Count
6,4,0,5,1,3,2,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0 6209
0,3,6,2,5,1,4,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0 6021
0,3,5,1,4,2,6,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0 1284
1,2,3,4,5,6,7,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0 678
1,2,3,4,5,6,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0 538
6,4,5,0,1,3,2,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0 518
First and last look like the Country Liberal ticket but with 518 voters swapping their 5th preference from the greens to the citizens electoral council - could be green-o-phobia or perhaps transcription error since the HTV does not show box C.....
Good stuff - happy to talk if you are interested....
Hi, Tim - thanks for this excellent example of using public data!
It inspired me to ask whether there's merit in a purely data driven approach to find and count the unique preference strings for each state - these can then be compared to the HTVs but will include non-HTV patterns such as donkey voting.
Trivial, fugly code and some findings at https://github.com/fubar2/aus_senate
I found '/' and '*' in the csv preference data - any idea what they are supposed to be? I just converted them to '0' to ignore...
For example, in the NT data, the top 6 duplicated patterns and their counts are:
First and last look like the Country Liberal ticket but with 518 voters swapping their 5th preference from the greens to the citizens electoral council - could be green-o-phobia or perhaps transcription error since the HTV does not show box C.....
Good stuff - happy to talk if you are interested....