Repository navigation
Improve Snippet function #399
Description
Activity
@ninimama If it's okay with you, I'd like to move this over to the discussion section as I think it would be healthy/useful to discuss/debate each of the points above. We should keep in mind that STUMPY doesn't try to be everything for everyone and our goal is to reproduce the published work.
Sure. I think it is good to move 1 & 3 to discussion. However, if you would like to reproduce the results of the paper, I think the snippet function should produce the indices of a set of subsequences similar to each snippet (e.g: check out Fig 24 of the paper or Fig 19).
However, if you think the ultimate goal of snippets is to merely provide a set of representative sequences rather than how they appear in the time series, then we can move 2 to discussion as well.
@ninimama I haven't looked into this part much so maybe you can help me understand how the indices are calculated? It appears that it has do with finding cross-over points.
However, if you think the ultimate goal of snippets is to merely provide a set of representative sequences rather than how they appear in the time series, then we can move 2 to discussion as well.
I think this is what I'm trying to get at with an open discussion. I am open to being convinced :)
This is part of the snippet code. If I understand correctly,
maskcontains the indices of distance profile corresponds to the snippet (in boolean). And, by summing over that, we calculate the fraction.So, if my understanding is correct, all we need to do is to return those indices.
Of course there might be some overlaps between one subsequence to the next, but we don't need to worry about that. We just need to return those indices. So, coloring them should give us the Fig. 24 of the paper.
Please feel free to close this one and move the questions to the discussion.
- locked and limited conversation to collaborators
on Jun 4, 2021

As I was working with snippets, I noticed the three following issues:
1- Apparently the indices of the snippet start exactly at the multiplies of m, which may not be the case for some data.
2- The function provides only one profile per snippet. However, that snippet repeats throughout the whole time series as higher fraction means it has been repeated more. So, it should be useful to have that information in the output, where I can see all the profiles (the starting index) of each snippet.
3- The fraction output is not sorted. I think it is better to sort it in descending order and accordingly provide the other outputs for the same order.