Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CHANGES.md
Original file line number Diff line number Diff line change
Expand Up @@ -92,6 +92,7 @@
* (Java) IcebergIO now writes rows containing `EnumerationType` (proto enum) fields as strings, instead of throwing `Unsupported Beam logical type Enum` ([#40299](https://github.com/apache/beam/issues/40299)).
* (Python) Fixed stateful DoFns with side inputs sometimes taking the timer key coder from a side input instead of the main input, which could make the worker fail to decode timer keys with `Unknown type tag` ([#40374](https://github.com/apache/beam/issues/40374)).
* (Python) `Duration` built from float seconds now rounds to the nearest microsecond instead of truncating, which could lose a microsecond ([#40263](https://github.com/apache/beam/issues/40263)).
* (Python) `ReadFromText` with `escapechar` no longer skips a delimiter that directly follows an escaped delimiter, which merged two records into one ([#40459](https://github.com/apache/beam/issues/40459)).
* Fixed X (Java/Python) ([#X](https://github.com/apache/beam/issues/X)).

## Security Fixes
Expand Down
2 changes: 1 addition & 1 deletion sdks/python/apache_beam/io/textio.py
Original file line number Diff line number Diff line change
Expand Up @@ -321,7 +321,7 @@ def _find_separator_bounds(self, file_to_read, read_buffer):
if self._escapechar is not None and self._is_escaped(read_buffer,
next_delim):
# Skip an escaped delimiter.
current_pos = next_delim + delimiter_len + 1
current_pos = next_delim + delimiter_len
continue
else:
# Found a delimiter. Accepting that as the next delimiter.
Expand Down
9 changes: 9 additions & 0 deletions sdks/python/apache_beam/io/textio_test.py
Original file line number Diff line number Diff line change
Expand Up @@ -1388,6 +1388,15 @@ def test_read_escaped_custom_delimiter(self):
self._run_read_test(
file_name, expected_data, delimiter=b'*|', escapechar=b'\\')

def test_read_escaped_delimiter_followed_by_delimiter(self):
with TempDir() as tempdir:
file_name = tempdir.create_temp_file(lines=[b'li\\\n\nne\n'])
self._run_read_test(file_name, ['li\\\n', 'ne'], escapechar=b'\\')

file_name = tempdir.create_temp_file(lines=[b'li\\*|*|ne*|'])
self._run_read_test(
file_name, ['li\\*|', 'ne'], delimiter=b'*|', escapechar=b'\\')

def test_read_escaped_lf_at_buffer_edge(self):
file_name, expected_data = write_data(3, eol=EOL.LF, line_value=b'line\\\n')
assert len(expected_data) == 3
Expand Down
Loading