-
Notifications
You must be signed in to change notification settings - Fork 15
Expand file tree
/
Copy pathNOTICE
More file actions
70 lines (48 loc) · 2.73 KB
/
Copy pathNOTICE
File metadata and controls
70 lines (48 loc) · 2.73 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
# Third-party notices
ColorSplitter is MIT licensed (see `LICENSE`). It contains code and model
definitions derived from the following projects. Attribution is kept here and
repeated in the file headers where the code lives.
---
## Resemblyzer — MIT
<https://github.com/resemble-ai/Resemblyzer>
Used for:
* `src/colorsplitter/core/audio.py` — the `preprocess_wav`, `normalize_volume`,
`trim_long_silences` and `wav_to_mel_spectrogram` implementations are inlined
from `resemblyzer/audio.py`. The librosa calls are kept argument-for-argument
identical so that embeddings stay compatible with existing checkpoints.
* `src/colorsplitter/core/hparams.py` — audio and model hyper-parameters.
* `src/colorsplitter/models/voice_encoder.py` — the `VoiceEncoder` network
definition, including `compute_partial_slices` and `embed_utterance`.
* The optional `speaker-upstream-v1` weight (`resemblyzer/pretrained.pt`).
The `Resemblyzer` *package* is deliberately not a dependency: only the parts
actually used are vendored, so the install does not pull in the whole project
and its transitive requirements.
---
## 3D-Speaker — Apache License 2.0
<https://github.com/alibaba-damo-academy/3D-Speaker>
Used for:
* `src/colorsplitter/core/cluster.py` — `SpectralCluster` and `UmapHdbscan`,
themselves adapted from SpeechBrain's spectral clustering implementation.
The clustering behaviour is preserved. `p_pruning` is vectorised and the probe
matrix is optionally sparsified for the eigensolver; `tests/test_cluster_equivalence.py`
pins the results against a verbatim copy of the original implementation.
---
## wav2vec2-large-robust-12-ft-emotion-msp-dim
<https://huggingface.co/audeering/wav2vec2-large-robust-12-ft-emotion-msp-dim>
Used for the `emotion` encoder. The model weights are downloaded on demand and
are not redistributed here; see the model card for their terms.
`src/colorsplitter/models/wav2vec2.py` reimplements the wav2vec2 trunk against
that checkpoint. The graph was written with reference to HuggingFace
Transformers' `modeling_wav2vec2.py` and
`feature_extraction_wav2vec2.py` (Apache License 2.0,
<https://github.com/huggingface/transformers>), in particular the
weight-normalisation convention and the input normalisation, so that the
numerics match. Transformers is **not** a runtime dependency.
`src/colorsplitter/models/safetensors.py` is an independent minimal reader for
the safetensors format (see <https://github.com/huggingface/safetensors>,
Apache License 2.0), written to avoid pulling in another package.
---
## Encoder checkpoints
The bundled training lineage descends from CorentinJ's Real-Time-Voice-Cloning
(<https://github.com/CorentinJ/Real-Time-Voice-Cloning>, MIT), specifically its
speaker-encoder training and GE2E objective.