I would split the metric into two bars: capacity used, and sources included. Capacity tells whether the window is crowded. Inclusion tells whether the right evidence was selected.