ECC in Intel core caches

By: dmcq (dmcq.delete@this.fano.co.uk),
Room: Moderated Discussions
Michael S (already5chosen.delete@this.yahoo.com) on October 19, 2016 5:31 am wrote:
> dmcq (dmcq.delete@this.fano.co.uk) on October 19, 2016 5:22 am wrote:
> > Gabriele Svelto (gabriele.svelto.delete@this.gmail.com) on October 19, 2016 5:04 am wrote:
> > > another anon (nothanks.delete@this.email.com) on October 19, 2016 2:55 am wrote:
> > > > That is exactly the link I used to come to the above conclusion. That core was released
> > > > 5 years ago and there are 3 generations newer Xeon now, hence my question.
> > >
> > > If you look at the datasheets for the v2 and v3 families they don't mention that
> > > feature which I assume means it was either removed or at least disabled.
> > >
> > > > It is interesting how little noise Intel made about this, and
> > > > how little information turns up with simple google searches.
> > >
> > > It's possible that Intel didn't find any customer interested in that particular feature.
> >
> > For many things nowadays it is enough to know something has gone wrong and
> > then just abandon a task and redo everything.
>
> And what if you have unrecoverable error in critical kernel data?
> BTW, is you kernel data structured in way that it knows which data is critical and which is not?
> I'd think, most kernel are unable to tell the difference.

It's quite usual to have a switch at the start of error handling to minimize the impact of recursive problems. One can normally tell if an area contains code or data and is part of the system or a task or is unused. On a system with only parity a while back I also had longitudinal codes to fix code and read only data if an error was detected by a scrubber which checked the memory periodically.

> > It all depends on how often
> > errors occur and will they be detected quickly or hang around the place.
> >
> > I wonder if anybody tries to detect cosmic ray showers or suchlike events as it is far
> > harder to detect single bit changes inside the processor itself. Really one has to do
> > something like triple modular redundancy to cut the possibility of an error down to
> > where one can stop worrying - not that that helped the first shuttle launch ;--)
>
> I think, the theory says that for typical "cosmic" error rates triple redundancy
> is less reliable than self checking with double redundancy.
< Previous Post in ThreadNext Post in Thread >
Thread (26 posts)
TopicPosted ByPosted
ECC in Intel core cachesanother anon
  ECC in Intel core cachesGabriele Svelto
    ECC in Intel core cachesanother anon
      ECC in Intel core cachesGabriele Svelto
        ECC in Intel core cachesdmcq
          ECC in Intel core cachesMichael S
            ECC in Intel core cachesdmcq
        ECC in Intel core cachesDavid Hess
          ECC in Intel core cachesanonymou5
            ECC in Intel core cachesDavid Hess
              ECC in Intel core cachesanonymou5
                ECC in Intel core cachesDavid Hess
                  ECC in Intel core cachesDavid Kanter
                    ECC in Intel core cachesdmcq
  ECC in Intel core cachesanonymou5
    ECC in Intel core cachesMichael S
  ECC in Intel core cachesDavid Kanter
    ECC in Intel core cachesanother anon
      ECC in Intel core cachesSimon Farnsworth
    ECC in Intel core cachesAaron Spink
    ECC in Intel core cachesDavid Hess
      ECC in Intel core cachesAaron Spink
        ECC in Intel core cachesDavid Hess
          ECC in Intel core cachesGabriele Svelto
            One way to handle uncorrectable errorsPaul A. Clayton
          ECC in Intel core cachesAaron Spink