Skip to content

[ENH] Feature Request: Introduce tsauditor.parsers sub-module for unstructured log ingestion #29

Description

@imann128

Most functions in tsauditor/profiler/, tsauditor/anomaly/, and tsauditor/leakage/ expect a cleanly structured tabular DataFrame. However, many real-world time series metrics start as unstructured text logs (e.g., application crashes, server syslogs).

What’s needed: Create a new sub-module tsauditor/parsers/ that handles unstructured text ingestion. This module will read raw log streams, parse out timestamps and categorical variables using regular expressions, aggregate occurrences into standard time frequencies, and return a fresh tabular DataFrame.

Why this is approachable: This keeps our core engine completely isolated. Instead of forcing our core analytics or remediation code to handle unstructured strings, this feature serves as a clean, non-destructive gateway that shapes messy logs into tabular frames before an audit ever starts.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requesthelp wantedExtra attention is needed

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions