Friday, November 27, 2015

T.O.C 27Nov2015

=====================   last updated: 17dec2015 ============================
  1. So you want to be a DSP architect?
    1. THE LONG STORY
      1. DSP Architecture today
      2. Why Matlab?
      3. The SSS (Sad State of Software) story
    2. BACKGROUND CHECKS 
      • Henneson and Patersy (CA)
      • Coprocessing
      • The "CORE"/the "CELL" / The "SLICE"
      • Benchmarking
        • BDT as an architectural tool
      • Fixed Point dialects
      • Algorithms and Data types
      • Platforms and SOC architecture
      • BUILDING BLOCK ISSUES
        1. Physical limits
        2. Design time, Build time and Run time
      • THE FLOW: METHODOLOGY AND TOOL ISSUES
        1. My stories
          1. Back to 1979: What is descriptive language? 
          2. Mapping vs Compiling: the never ending story?
          3. Soft vs Hard Macros - is Firm the answer?
        2. NPU lessons
        3. Retargetable compilers
        4. Simulation: we KNOW profiling is THE key! got the answer?
        5. Beserkley or Berkeley? and Alberto during that time...
      • FOUND IN THE WEBS
        1. The Matlab Engine 
        2. Is Kurt Keutzer the Donald Trump of Hardwired Processing? 
        3. Another Berkeley Randy Cat?
        4. Found in the cobwebs: my garage
        5. Jeff Bier is promoting CNN! 
        6. Andre Dehon is promoting Multics!
        7. The Trailing Edge
      • DSP (uP) (HISTORY OF ~ ARCHITECTURES) 
        1. DSP of the First Kind (1980-95)
        2. DSP of the Second Kind (1995-2005)
        3. DSP of the Third Kind (2010- ?)
        4. DSP of the lost kind (1995-2005)
        5. DSP  of any kind (1950-2050)
        6. DSP vs Micros
        7. DSP vs FPGA
        8. The Archives and the DSP historian
      • FROM DSP TO DSP
        1. This wonderful world of DSP
          1. A world? More like a sect! 
            1. The Pope (Will), the Cardinal (Jeff) and the Wizard (Gene)  
        2. The DSP Old Timer's Club.
        3. Stop me if you've heard this one before!
        4. DSP history
          1. Bit slice: a saga 1975-1985
          2. Building Block:  was 1992 and IDT  the last DSP BB? or Weitek?
      • TARGET BBs
        1. Analog programs (Basic, HPC and other gizmos)
        2. TI C25 Tips and Tricks 
        3. SPmag Tips and Tricks
        4. SPMag 1990-2005
    3. BB Level 1
      1. Arithmetic BB
        1. add,mul,compare, conditionals, degenerated, compound, muladd 
      2. Bit wise BB
        1. Shift, 
        2. Bit-field, 
        3. Byte field, 
        4. Bit count logic (cls,clz, clo. etc..) 
        5. Boolean processing 
      3. Vector BB
        1. Vector arithmetic, 
        2. Memory vector search, structured tables, lookup
        3. Shapes (Matlab)
          1. Shape generation
          2. Matlab primitives
          3. Reporting
    4. BB1 DSP STRUCTURES
      1. DSP trilogy 
        1. ALU
        2. BMU
        3. MULT  
      2. MAC
      3. DAU
      4. AGU 
      5. PCU
      6. RF (Register File) and MEM
      7. VPU (Vector Processing Units)
      8. SU (The Shuffle Unit) 
      9. Bit slice: Bit slice building blocks
    5. BB2 DSP FUNCTIONS
      1. Filter
      2. FFT
      3. Correlators & bit comm. engines
      4. Sampling functions
      5. Matlab dsp functions
        1. accumarray
        2. diff
        3. sum
    6. BB3 MATLAB CONSTRUCTS
      1. Find
      2. Predication with Matlab
      3. Matrix functions
    7. BB4 MATH 
      1. Divider
      2. Math.h
      3. Complex numbers
      4. Further with numerical recipes
    8. BB5 CODING
      1. Basic Coders
        1. Coding generalized engine
        2. Gray, Hamming, etc..
        3. Parity
        4. Cyclic codes: Firecode
        5. CRC
        6. Huffman coding
        7. Arithmetic coding
        8. Gallois Fields
      2. Convolutional coders
        1. Gsm decode
        2. Hard decision
        3. distab
        4. distaabcc
        5. Viterbi equalizer
        6. Viterbi butterfly
      3. Com. Block coders
      4. Iterative coders
        1. Turbo encoding
        2. Turbo decoding
        3. LDPC
      5. Walsh, Hadamard and cdma
        1. FHT
        2. Happy Chirper.
    9. 5 MATLAB TOOLBOXES
      1. NUFFT
      2. SAR
    10. APPENDIX 

    Sunday, November 22, 2015

    Is Andre Dehon a genuine philosopher and Multics misunderstood?

    The answer is: YES
    =====================================================================
    That's beat all for this week!

    We all know how bad Multics was and so Unix was created. From that point history has been written:

    • Unix is the white knight....mac OS ...embedded linux
    • Microsoft windows/dos/PC  is the dark side
    • Multics? dead and buried.. never again.

    So, imagine my surprise, while checking some older MIT stuff from AD (Andre Dehon) I came across the Multics Reunion  MIT 2014 with a paper of AD.

    Abstract: At a time when computers are increasingly involved in all aspects of our lives, our computer systems are too easily broken or subverted.  The current state of affairs is, no doubt, unsurprising to Multicians who are painfully aware of the design and security compromises that went into the base design of today's mainstream systems.  The past 30 years has also brought vast changes in the availability and costs of computer hardware as well as significant advances in formal methods.  How do we exploit these advances to make computer systems worthy of the trust we are now placing in them?  We specifically take a clean-slate approach to computer architectures and system designs based on modern costs and threats.  We spend now cheap hardware to reduce or eliminate traditional security-performance tradeoffs and to provide stronger hardware safety and security interlocks that prevent gross security and safety violations even when there are bugs in the code.  We embrace well-known security principles of least and separate privileges and complete mediation of operations.  Our system revisits many pioneering Multics concepts including gates between software components with different-privileges, small and verified system components, and formal information flow properties and guarantees.
    Project paper: http://www.crash-safe.org     

     Now, I consider AD a madman since he went back to the east coast instead of staying in sunny California. But then AD was(is) always the ultimate contrarian, lateral thinker, deep searcher of truth, and so I should not be surprised by such a support on Multics..
    Ten years ago he was fighting for the replacement of silicon platforms by {bio? to be checked}. And asked about the prospect, he rightly said that he was an academic and would not dare fighting the semiconductor industry...

    Is there a lesson to learn?   Obviously, cutting corners and bowing to economics pressure will bite you in the long run. Was Multics the answer to hacking? I have my doubts (but deep interest in this kind of rhetorical questions).

    And I cannot not leave without mentioning Andre Dehon bio
    Andre DeHon received S.B., S.M., and Ph.D. degrees in Electrical Engineering and Computer Science from the Massachusetts Institute of Technology in 1990, 1993, and 1996 respectively.  From 1996 to 1999, Andre co-ran the BRASS group in the Computer Science Department at the University of California at Berkeley.  From 1999 to 2006, he was an Assistant Professor of Computer Science at the California Institute of Technology. 
    In 2006 he joined the Electrical and Systems Engineering Department at the University of Pennsylvania, where he is now a Full Professor.  He is broadly interested in how we physically implement computations from substrates, including VLSI and molecular electronics, up through architecture, CAD, and programming models.  
    He places special emphasis on spatial programmable architectures (e.g. FPGAs) and interconnect design and optimization.  

    Multics BIO:  Andre DeHon is a bastard child of the tail end of LISP Machine and Multics eras, having been a research assistant for Knight and a teaching assistant for Saltzer.  As a member of MIT's Student Information Processing Board (SIPB), he was part of the group that pushed Multics access to MIT students and was logged in during the decommissioning of MIT-Multics. So, while he never contributed to Multics, he was around in time to learn that there were computer systems that predated Unix and Windows and that did have a principled way to address safety and security.  He hopes the world is now ready for many of the Multics and LISPM ideas that were ahead of their time and have mostly been forgotten during the dark ages of mainstream Internet growth.

    1.4 THE FLOW: METHODOLOGY AND TOOL ISSUES

    The goal of this section is to put together all topics relevant to methodology and tools

    The proposed M2IMPIR flow

    M2IMPIR (Matlab2Implementation) quick description 
    1. Matlab BB (Building Blocks)
      1. Develop specific BBs. 
      2. Reuse BBs, out of one of the several thousands of Toolboxes and librairies.
      3. Build AS-DSP and simulate it.
    2. 2 Map to Ideal Platform (BB to BB)
    3. IMPlementation with Ideal Resources
      1. Run and verify 
      2. Iterate 
    M2IMPIR issues
    • There are 3 types of issues: Matlab, Mapping, Implementation platform. 
    1. Matlab: Here the problem is "Matlab BBs". 
      1. Because, more often than none, a Matlab BB ends up as a FPGA IP, i.e hdl code, i.e concurency. And Matlab is everything but concurent. 
      2. The other issues have to do with the language. Going through all of them is outside this section just say the number is large and many are not obvious.    
    2. Mapping:  Mapping has a few pitfalls but they are well known. 
    3. Implementation platform:
      1. My current methodology is to use a platform with ideal resources (which is not the same as infinite resources). That's all. 
      2. Even an ideal platform has some issues 
        1. but at least the flow is managable. 
      3. For those  interested in implementation platforms and architecture, I propose to start with  Andre Dehon web site (at Uni Penn) , read all papers and go from there.

    Further Topics

    1. My early life: looking for an oxymoron: the descriptive Language
    2. The NPU lessons
      1. The thesis of Norwich Rich. 
    3. Retargetable compiler: from bit slice to Archelon.
    4. Meet the infinite loop: simulation and profiling on RC platforms (reconfigurable computing).

    Friday, November 20, 2015

    Jeff Bier is promoting CNN

    ----  Jeff Bier is promoting CNN! ------

    Thesis: Neural Networks? not again!

    I used to say " when I hear the word RISC (or speech recog) I draw my guns". Things have slightly changed so instead theses days, my current one is "when I hear Neural Network ..."
    I never had a really good fight with NN because I never met NN in the open. 
    NN was an esoteric digital technique (like many others) found in the DSP conference papers. Over the years I have kept NN implementation of DCT, FFT, MAC, etc.. and never had a chance to take a look.
    But recently I have seen  many micro-controllers promoting the technique. 

    Antithesis: Neural Networks are alive and well

    1. Matlab has a Toolbox
      1. and maybe much more than that. 
    2.  Jeff Bier put his money where his mouth is:
      1. The Embedded Vision Alliance invites you to attend an exciting event --> see below:
    Synthesis: No thanks, I pass on that one. I still think it is bleeding edge. 


    Wednesday, November 18, 2015

    DSP architecture today


    The goal of this section is to put in one place any architecture bits and pieces which are interesting for this blog. 

    The generic design methods                                                                                                  
    • Vector processing 
    • SIMD, swp (sub-word parallelism), multi lanes, packed arithmetic... 
      • where to stop? 256? 1024?
      • why stop at 1K? Matlab row vectors, cln vectors, arrays are a magnitude higher 
      • Minimum parallelism= 64K 
        • simd of 16  x 16 clusters x 16 cores  x 16 modules 
      • Golden ratio (x4, x16, x64, x256) 

    On My watch                                                                                                                          

    • "ASPI world"
      • Andre Duhon
      • FCCM
      • Xilinx at Large
    • GPU
      • Matlab and GPU
      • Coda
      • Deep packet searching
    • Hot Chips
    • TenSilica cores and domain specific platforms, 
      • Cadence

    They still Matters                                                                                                                  

    • ARM Cortex A - family and evolution
      • how far will they go with speculation? 
      • repeating the same mistakes in a new way?
    • Intel ecosystem
      • Altera integration
    • TI Platforms evolution 
      • replacing DSP core with hardwired COP

    Lessons Learnt                                                                                                                         

      • NPU

      ---------------                                                                                                                               


          Saturday, November 14, 2015

          SPmag1990 jan

          SPmag 1990 jan

          (this post was written while listening to: Dead Europe 72 interleaved with NRPS first album)

          We start lucky. In this issue, the main article is a survey of multiplier techniques by the master himself (Fred J. Taylor). Also interesting is the introduction of Matlab version 3.5, including the all new signal processing toolbox. All the book reviewed and many items are relevant to this work. We will NOT finish with a very fashionable topic: Neural Network.  

          BB struct                                                                                                                        

          The multiplier (fred J. taylor)

          This is the DSP BB by excellence. There is hardly any algo which does not rely on a multiplier and it is not cheap. As Fred mentionned, the 16x16 multiplier occupies 40% of the 32020 chip. And I believe 80% of non-memory real estate.  
          Now, in our world of ideal resources why do we bother with the size of the Mult? Because ideal does not mean infinite. If we develop (say) a 2000 Million node chip, each node would better be optimised. And the best choice for a node (or PE) is a MULT. The next best choice is a DSP core, where MULT is still 80% of the core. (it reminds me of the debate at Xilinx to replace the embedded 18x18MULT with a C25 like core to extend their market share; all people cannot be right all of the time)
          So here are the techniques (flattened) given by Fred:
          • Shit add or iterative
          • Booth algorithm (and modified~)
          • Wallaces trees
          • Cellular arrays ( Perazis,  Baugh-Wooley)
          • Systolic Array Multiplier
          • Bit Serial
          • Distributed arithmetic
          • Canonic signed digit number systems
          • Logarithmic number systems
          • Residue Number sytems  
          With the hindsight of 25 years, what to think? Well it stands pretty good, except for the systolic array.
          What about the non-traditional number systems? Frankly we are very NOT positive about these, because the problem is the time lost in the interface. 
          As for the others, bit serial and DA are common FPGA techniques and are available in Matlab too  (how to generate HDL code for a lowpass FIR filter with Distributed Arithmetic (DA) architecture). 
          The parallel techniques are standard logic techniques but they are still the most promising because the exploration space is very vast.
          We have a favorite, which is a kind of subword parallelism. The longest delay in a 16x16 MULT is the final 32-bit adder. This delay can be halved by splitting the adder in 3 16-bit adders. By the same token we can subdivide further with byte and nibble adders hence reducing the delay from 30 to 6 logic gates. We simulated this type of architecture in Matlab (using FP!) (see elsewhere).    

          Workshop: VLSI for Signal processing at Iccasp90 (Edward Lee)

          It is mentionned that several Experimental Parallel Machines for signal processing will be explained. OK let us see.
          With the hindsight of 25 years, what to think? A bit of disaster array (sorry area) and we know the results (the GPU). However, in the M2IMP world where the resources are ideal, and the software is straightforward, each of these EPM will be given a good look in due time.

          Vector Quantizer (Tran & all,  IEEE tr. on com, sep 89)

          It is worth asking if the VQ is a structural or functional (speech specific) BB. 
          In our experience we consider it as a generic unit with app. specific parameters. 
          Matlab?  It is an object in DSP system toolbox .
          With the hindsight of 25 years, what to think? This is a good example where Matlab is now the reference. 
            

          BB DSP fun                                                                                                                     

          FFT workshop at ICCASP90 (John Cooley) 

          In this workshop the FFT master explains all the options and innards, including the different radices. Great!
          Matlab?  FFT is a built-in. The number of samples is totally flexible. The problem is that it is a black box. Not quite all dark as it is based on FFTW. But we had a mixed experience with the golden model methodology. While trying to develop a radix 9 FFT and comparing it to Matlab, stage by stage. As a matter of fact it was easier in Excel. As I am writing this, it does not make sense since in effect I am saying that FFT is non determinstic. But so is my memory of it. I do remember Marc having a lot of trouble with Viterbi decoder but it is different cause because it is a treillis.

          Further Transforms - workshop at ICCASP90 (John Cooley)


          In this workshop, John goes further with number-theoretic transforms(NTT) and polynomial transforms. 
          Matlab? As far as I know they are not included. Fortunatly the community supplies them 
          • NTT: NTT.m on matlab central. 
          • PT brought me to the site of chemnitz, Daniel Pots. In turn brough me the NUFFT toolbox. 
          With the hindsight of 25 years, what to think? 
          • FFT: is now old hand, {see elsewhere}. 
          • NTT; I have no clue, but a NTT engine would be nice. 
          • NUFFT (non uniform FFT) an extremely promising field which I bundled with CS (compressed sampling).  

          Communication paper: dpll

          Matlab? not included; see ecosystem:  Modeling and Simulating an All-Digital PhaseLocked Loop By Russell Mohn, Epoch Microelectronics Inc.  

          Communication paper:  sigma delta 

          Matlab? included

          BB in DSP domain                                                                                                           

          A tutorial at ICASSP 90: SAR (Synthetic Aperture radar) (David Munson) 

          Matlab?  Included, excellent Block diagram and tutorial.  


          BB in Matlab specific                                                                                                      


          These Building Blocks are Matlab specific because if they exist it will be more likely in the shape of M-code than anything else. Never met them in DSP software libraries or commercial chips. 

          Signal processing Toolbox is introduced {advert for Matlab3.5}
          First ever SP toolbox. This includes a large number of functions many of which are not very common. 

          Non Linear Spectral Estimation Methods (Prabhu &all, IEEE proceedings, June 1989)
          Here is an article which covers algos than a standard DSP cannot implement. Among these techniques are arma, music and esprit. 
          Matlab? They can be found over diverse Matlab toolboxes.   
          Also, many of them can be found in the very professional looking HOSA toolbox which is available from the Central.  {HOSA - Higher Order Spectral Analysis Toolbox by  12 Feb 2003 (Updated Spectral and polyspectral analysis, and time-frequency distributions.}

          SUMMARY TABLE 




          Monday, November 9, 2015

          The DSP Historian

          DSPs - Archives

          1. The DSP Historian (see below)
          2. 1978 - 1987: the call-Up
          3. 1988 - 1995
          4. 1996 - 2004
          5. 2004 - Today 
          6. Pre History

          The DSP Historian:
          The goal of this section is to go through all the papers having for topic "the DSP uP chips". This will become clearer as we go in order of things.

          Background
          The knowledge of " History of DSP uP chips" is a major tool in my "alternative world" methodology to develop AS-DSP IP. As explained elsewhere I believe that anything after 1990 has little value compared to whatever happens before. Hence the "History of DSP uP chips" become 2 major fields:
          • before 1990: the chips ans architecture to know 
          • after 1990: the field of ideas: sounding board, the mistakes, the real advances, etc..
          Taking that at face value, there are plenty of buddying architects who  would like a quick summary of DSP uP chips before investing further time.
          Hence, the issue becomes:  where is the best story? "who is the best DSP historian?"

           Anything in my garage?                                                                        
          I will do the usual and go through my stash. In one folder I found stuff which should be available on the web, but not always free.
          • Noted, in a 10 page paper [1], 
            • the first generation dsp (amd2900 and TI C10) followed by the 2nd generation (C25, dsp16). This is so wrong that you have to be admirative of the author(a lateral thinker? an hyperspace traveller?).  
            • Anyway, the point here is that by studying the history of DSPs, the architect is obliged to classify (by generation, by features) and hence make choice, which as we all know, the lifeline of a DSP architect. If you cannot even classify DSPs don't bother.
            • Also noted are references to the pipeline, including the mind-boggling <data stationary versus time stationary>.  This, I believe, is a complete red herring but in the world of free world DSP, anything is fair and square.
            • OOh, and the DSP32C has a reservation table!! 
            • Finally, another "gem" is Edward Lee's "interleaved pipelined" architecture. Hum?
          • Speaking about Edward Lee, not only he is the grand father of many things (see BDT) but also the father of the classical paper[2,3], a first, dated 1988 < programmable DSP architectures>. If you have access to the paper, you can stop here. 
            • There is an IEEE micro version dated 1990 [4], in which Lee conservativly dropped  the architecture term. 
              • In these days, only a few people could risk using the term.(see MPR oct 1989)  
            • In my case, Lee was was far from being the first one. Since 1978, I had already written 4 different versions of History of DSP chips including a 5m long mural, so Lee was of little value.
          • At this stage it is worth mentioning the TRILOGY (Will, Gene, Jeff). What where they doing in 1988?For different reasons, none of them had written"THE REFERENCE". 
            • Jeff: obviously, being a student of Lee, postdate his work. 
            • Gene: being in full war with his competitors, he was not in the position to write objective reports. Still, good grunting noise could be heard at TI conferences.
            • Will: I came across his first(?) report circa 1986. Primarily a marketing document, but in these days it was a technical gem.    
          • Following this, here are much more recent papers, freely available, all being university courses on DSP . 
            • With the title, Special Purpose Digital Processors (DSP) [5], a new type of acronym is introduced: the warped acronym. Why not? e.g. General Purpose Processor (CPU). It is a rather exhaustive 29 pages lecture note document from TUHH.
            • Less exhaustive a Texas loaded ppt from Austin[6]
            • A french in french 144 slides pdf from Irisa [7]. Note the T.O.C is perfect for a classical story of DSPs.
              • I. Introduction
              • II. Architectures MAC/Harvard
              • III. Evolutions des DSP
              • IV. Flot de développement (Flow) 
            • Finally [8] is a french paper in English from 2002. The authors are from McGill.
          Gene's law
          As already mentionned Gene Frantz was kind of busy proseletyzing the world to the DSP mantra till in 2000 he was convinced to write a Millenium paper [9] for IEEE micro in which developped his popular (among us DSP architects) Gene's law. Gene's law is the equivalent of Moore's Law.

          First for the not so good.
          • Here is the lesson for us all architects: we DON'T KNOW HOW TO limit ourselves to history. It is always past, present, future. Extrapolation is what sells the paper and 9 out of 10 we are wrong. 
          • Unfortunatly this paper was badly timed as Moore's law (speed) was going broke.. 
          • Gene extrapolated for the 2010 and came with the 10GHz DSP . 
            • Myself, around that time I had papers extrapolating speed for CPU and embedded CPUs such that 2010:10G __ 2020: 20G__ 2030:30G for CPUs you see the trend. Embedded CPUs had a lower slope and were around 5G in 2010.
          • Now, even better is the Moore analogy. Little known but true, in multiple interviews circa 1980 Moore explained that with ultra VLSI, except for FFT he had no idea how to fill a logic chip. Good one, Gordon, no wonder you came up with the !IAPX432. In the same way, is Gene, turning the TI slogan on its head "limits to your imagination" here is the excerpt    (Determining how to use that processing power effectively will require imagination that goes beyond conventional engineering methodologies.)...ooooh my oh my!
          • In all fairness, the paper was called DSP trends, so...
          Now, for the good things, .
          • Moore's law is broken but I would not be so sure about Gene's Law
          • Gene changed architecture by putting power consumption (and efficiency) to the forefront.
            • This is what matters to me. Effectivelly it means than in the end, after all low hanging fruits are gone,  there is only 1 technique left to improve efficiency: hardwired is replacing software. Customization is replacing cpu clock.{at this stage parallelism and multicore is just an implementation detail} 
          • Gene extrapolated the 50 GIPS DSP which did happen and even by today seems pretty tame. 
          • Gene mentionned a lot of stuff, details that I picked up and in the end, this is it. MY MOST RECENT PAPER ON THIS TOPIC WITH ANY VALUE. 
            From Lee to BDT
            By the 1990s, Jeff Bier was starting an exceptional business <www.BDTI.com> B standing for Berkeley (see the story of B&B elsewhere). As part of his duties to the DSP community at large, Jeff was given lectures on DSP chips. So if you've read Lee's seminal paper the next logical step is a Jeff's paper. Which one? No idea, I have 41 in my stash. Even Jeff does not know how many he made so better go the BDTI web site.
            Also, there is much more to BDTI than DSP chips history and trends. They are THE reference in DSP benchmark (see elsewhere).   

            Krishna Yargalada
            By 1996, BDT had the monopole of "DSP chips - history and trends". Serious talk about breaking it, was heard in congress. Courageously Krishna Yargalada attempted to infiltrate DSP conferences with his own version[10] but t,hinking about it, it was a clever disguise to launch Hellosoft in what became the big revolution of our world of DSP "outsourcing assembly to India" {see elsewhere}.

            Henry Davis
            Hi Henry! 

            Robert Cushman and the EDN DSP directory
            From 1981 to 1988(?), Robert Cushman, a major contributor to EDN,  wrote many articles on DSP. He put together the first EDN (microprocessor-like) DSP directory in 1987.  I have all of them till 2008. Going through them in order, gives an exceptional view of history.

            THE UBER REFERENCE
            Noted with pleasure that a lot of papers put as reference the BDTI classical book.
             Phil LapsleyJeff BierAmit ShohamEdward A. Lee , “DSP Processor Fundamentals: Architectures and Features,” IEEE Press, 1996.
            This book is THE BIBLE reference for DSP architect (see elsewhere). It is much much more than an introduction to DSP chips.

            References                                                                                                                                      
            1. <honestly I dont know; lost in history>
            2. Edward Lee "Programmable DSP Architectures Part 1" IEEE ASSP oct1988
            3. Edward Lee "Programmable DSP Architectures Part 2" IEEE ASSP jan1989
            4. Edward Lee "Programmable DSPs - a brief overview" IEEE Micro oct1990. This is essentially the intro of paper 1 + an excellent  summary table.  
            5. F. Mayer-Lindenberg, "Special Purpose Digital Processors (DSP)" TUHH (Hamburg University) lecture notes, dated ? 
            6. Brian L, Evans "Introduction to DSP" University of Texas, Austin
            7. <http://www.irisa.fr/R2D2>   Olivier Sentieys, IRISA, 2005
            8. Benoit Champagne, fabrice Labreau "An introduction to Digital Signal Processors" Compiled 2002 
            9. Gene Frantz "DIGITAL SIGNAL PROCESSOR TRENDS", IEEE micro nov-dec 2000 p.52
            10. Krishna Yagarlada "DSP chips for communications"  which conf?? 2003?? (see excellent slide)
            This is technical slide, illustrating different architecture choices !! 
            Editor notes: no links are given as they become obsolete. Google keywords instead!

            Saturday, November 7, 2015

            The philosophy of BB

            The philosophy of BB

            1. A BB is either simple or complex. It should never be complicated. 
              1. The difference between complex and complicated is that complex can always be broken down in a sum of simple things. 
              2. And complicated is well .. complicated. For instance a lot of software is complicated.
            2. People familiar with the evolution of  electronics, remember the SSI, MSI, LSI stages and can put the heydey of BBs as the days of Bit Slice. 
              1. A bit slice was built with simple elements and a bit slice was itself the building block used in upper dsp Building Blocks, dsp functional Blocks (FFTer, Filter, Correlator), dsp CPU  or even the DEC mini-computer. 
            3. For instance a standard dsp CPU (called DSP thanks to TI) are made of 4 standard bit slice blocks (ALU, MULT, AGU,  LSU) + memories + I/O Space.
            4. A very good DSP design should have a maximum of 3 hierarchical levels (MSI, LSI, Final product). 
              1. Anything above 5 levels is prohibitive and should not be tackled here.
            5. Because we rely on a hierarchy of blocks, testing of each block is vital to the whole process. 
              1. A bug in a low level block is a disaster since it will be repeated hundred of times over several final products. 
            6. Same thing for an inefficient block but with a twist.
              1. Contrary to a bug, efficiency is a relative term. 
                1. You can redesign the BB with a different name. 
                2. (for sophisticated designer) you can redesign an upward compatible block  with the same name 
                  • but then you need a very hefty test suite.
            7. So what about repairing a bug? 
              1. In my book, a bug is both visible and NOT compatible with itself. 
              2. Hence if you repair a bug you don't know the result; somewhere down the line somebody wrote software taking this bug into account or somebody used the BB taking the bug into account.
            8. This is the INTRINSIC problem to the BB approach. You are going BOTTOM UP in the definition. So forget about reparing bugs. Instead create another BB.

            Matching functions to structures

            ====== Building Block ISSUES  ======= 

            Matching functions to structures

            When designing a new BB it is because you need it. Hence, we call this need: a strictly functional need. There are two ways to go:
            - stick to it and solve the present issue.
            - have a broader view and design a more general purpose brick which can be reused.
            A third case happens when the functional block can be turned into a generic structural block for no additional price. So before designing a new block, it is worth asking oneself.
            Can I make this BB more generic?
            Let us take as example the addition needed in a complex multiplication:
            zr = xr*kr -xi*ki
            zi = xi*kr + xi*kr
            same as
            zr = p1 - p2
            zi=  p3 + p4
            If we use a dual multiplier, the 4 multiplications can be done in 2 cycles.

            Then it must be followed by two BBs for the complete equations. We now write the equations as follows.
            1.   [zr zi] = thru22(p1,p3)   %   zr=p1 zi=p3
            2.   [zi  zr]= addsub42(zi,p4,zr,p2) %zi= zi+p4  zr=zr+p2
            It is easy to see that the two BBs are generic but the inputs and the equations are not straightforward..

            The standard alternative is
            1. zr = sub(p1,p2)
            2. zi=  add(p3,p4)
            which is simpler. So what are taking about here?

            In fact there are two issues with the standard way;
            1. It needs the 4 input data on the first dual multiplication
            2. It does not work when we want  real_imag  to be in the same register so that the complex number is considered an entity made of 2 parts..

            Thursday, November 5, 2015

            T.O.C 4Nov2015

            1. So you want to be a DSP architect?
              1. THE LONG STORY
                1. DSP Architecture today
              2. BACKGROUND CHECKS 
              3. BUILDING BLOCK ISSUES
                1. Physical limits
                2. Design time, Build time and Run time
              4. METHODOLOGY AND TOOL ISSUES
              5. FOUND IN THE WEBS
                1. The Matlab Engine 
                2. Is Kurt Keutzer the Donald Trump of Hardwired Processing? 
                3. Found in the cobwebs: my garage
                4. The Trailing Edge
              6. DSP (GP) ARCHITECTURES 
                1. DSP of the First Kind (1980-95)
                2. DSP of the Second Kind (1995-2005)
                3. DSP of the Third Kind (2010- ?)
                4. DSP of the lost kind (1995-2005)
                5. DSP  of any kind (1950-2050)
              7. FROM DSP TO DSP
                1. This wonderful world of DSP
                  1. A world? More like a sect! 
                  2. The Pope (Will), the Cardinal (Jeff) and the Wizzard (Gene)  
                2. The DSP Old Timer's Club.
                3. Stop me if you've heard this one before!
                4. DSP history
                  1. Bit slice: a saga 1975-1985
                  2. Building Block:  was 1992 and IDT  the last DSP BB? or Weitek?
              8. EXAMPLES of Target AS-DSP
                1. Analog programs (Basic, HPC and other gizmos)
                2. TI C25 Tips and Tricks 
                3. SPmag Tips and Tricks

            2. BB Level 1
              1. Arithmetic BB
              2. Bit wise BB
              3. Vector BB
            3. BB1 DSP STRUCTURES
              1. Processing trilogy (ALU, BMU, MULT)  
              2. MAC
              3. DAU
              4. AGU 
              5. PCU
              6. RF (Register File) and MEM
              7. VPU (Vector Processing Units)
              8. SU (The Shuffle Unit) 
              9. Bit slice: Bit slice building blocks
            4. BB2 DSP FUNCTIONS
              1. Filter
              2. FFT
              3. Correlators & bit comm. engines
              4. Sampling
              5. Matlab functions
            5. BB3 MATLAB BB
            6. BB4 MATH 
              1. including Complex numbers
            7. APPENDIX 

            Is Kurt Keutzer the Donald Trump of Hardwired Processing?

            The answer is No!

            The link
            http://www.eecs.berkeley.edu/~keutzer/

            Background
            In the sputtering days of  CPU Architecture, just after the millenium, I remember KK (and Berkeley in general) as a different voice as to which directions to take for future CPUs. Effectivelly here are a few arguments of the time (*):
            -  we are now in a hardware cycle as opposed to a software cycle
            -  the next step in MIPs is NOT reached by increasing the clock
            -  instead programmable logic or FPGA or hardware or ASIP (Application Specific IP)
            -  hurray for configurable core (typ. tensilica,  then stretch)
            -  hurray hooray for reconfigurable computing (all dead).

            Jump to 2015. As I am going through my notes on possible architecture trends (of interest to this Blog), I have 2 KK papers that I want in electronics format. Google the paper name + author and got nothing. Lost in the ozone. Okay I will scan them.
            To cut a long story short, the KK I found is not the one I had in mind(*), but still an exciting guy to follow..
            Anyway, here is  his web page; among other things, KK mentions that he lost money with Catalytic, sorry, I contributed to that Kurt!
            Here is his current interest :
             Exploring Design Patterns for Parallel Computing
            which is somehow related to this Blog.
            I downloaded 2 papers:
            - how to map recursivity in hardware
            - speculation.
            And the other papers are good too!

            (*) as I remember; I might be confused, not sure if KK was leading all that;

            1.1 The Long Story


            1.1 - THE LONG STORY


            The goal of this section is to give some meat (and the background) to our base statements.But it does not try to discuss or justify any of our choices.
            • As mentionned in the introduction, we are in the very comfortable position of presenting a system which is outside reality. 
              • In effect, all decisions have NO tangible reasons (such as based on cost, design time or available tools and platforms).
              • It is all in the eyes of the beholder.

            Q&A


            • DSP architect is Dead?
              • Today DSP architect job's split: 95% platform, 2% Core, 3% tools.
              • PLATFORM: 
                • Is there a DSP architect in the house? No, he is split into pieces
                • The system guy (m-code and simulation) --> 90% of the work
                • The DSP guy (*) --> 10% of the work
                • The firmware guy: can take the job of the DSP guy, especially writing in C
                • The ASIC/FPGA designer:  can take the job of the DSP guy.
            • From above, the REAL DSP architect is the system guy? yes
              • But this guy is not (does not want to be) an architect
              • {..} Here would come all the efforts such as Matlab to C/verilog, System C, etc,, and in general all the failures of the Tool vendors to provide a flow:system to implementation.  


            • Why customisation?
              • Because we used to be vertical, went horizontal and now we are back to vertical.
            • What is accelerated processing? 
              • see  IEEE micro july/aug 2008
            • Why Matlab? 
              • Because it is the most popular system tool.
              • Also Matlab is much more than a language (randy's 3 things 1):
                • a language, *** 
                • an EDE  ****
                • a simulation environment *****
              • Matlab is not a concurent language (contrary to verilog). Does that prevent you from sleeping? here i avoid the use of the word parallelism because of the funny parfor. 
            • What is mapping?
              • An analogy: mapping is hardmacro, compiling is soft macro
              • An approximation: mapping is by hand, compiling use a tool.
              • An aromatism: mapping is an assembler macro, compiling is C
            • Can mapping be used with a compiler? 
              • You bet! As a matter of fact, it is the only way to be succesful.

            Monday, October 19, 2015

            T.O.C November 2015

            1. So you want to be a DSP architect?
              1. THE LONG STORY
                1. DSP Architecture today
              2. BACKGROUND CHECKS 
              3. BUILDING BLOCK ISSUES
                1. Physical limits
                2. Design time, Build time and Run time
              4. METHODOLOGY AND TOOL ISSUES
              5. FOUND IN THE WEBS
                1. The Matlab Engine 
                2. Found in the cobwebs: my garage
                3. The Trailing Edge
              6. DSP (GP) ARCHITECTURES 
                1. DSP of the First Kind (1980-95)
                2. DSP of the Second Kind (1995-2005)
                3. DSP of the Third Kind (2010- ?)
                4. DSP of the lost kind (1995-2005)
                5. DSP  of any kind (1950-2050)
              7. FROM DSP TO DSP
                1. This wonderful world of DSP
                  1. A world? More like a sect! 
                  2. The Pope (Will), the Cardinal (Jeff) and the Wizzard (Gene)  
                2. The DSP Old Timer's Club.
                3. Stop me if you've heard this one before!
                4. DSP history
                  1. Bit slice: a saga 1975-1985
                  2. Building Block:  was 1992 and IDT  the last DSP BB? or Weitek?
              8. EXAMPLES of Target AS-DSP
                1. Analog programs (Basic, HPC and other gizmos)
                2. TI C25 Tips and Tricks 
                3. SPmag Tips and Tricks

            2. BB Level 1
              1. Arithmetic BB
              2. Bit wise BB
              3. Vector BB
            3. BB1 DSP STRUCTURES
              1. Processing trilogy (ALU, BMU, MULT)  
              2. MAC
              3. DAU
              4. AGU 
              5. PCU
              6. RF (Register File) and MEM
              7. VPU (Vector Processing Units)
              8. SU (The Shuffle Unit) 
              9. Bit slice: Bit slice building blocks
            4. BB2 DSP FUNCTIONS
              1. Filter
              2. FFT
              3. Correlators & bit comm. engines
              4. Sampling
              5. Matlab functions
            5. BB3 MATLAB BB
            6. BB4 MATH 
              1. including Complex numbers
            7. APPENDIX 

            The Matlab Engine

            The Matlab Engine

            This is a piece of IP which is free and can be readily implemented in a fabric or FPGA. Even better, it could be already included; the same way than a Cortex core is included in a high end Altera FPGA.
            Its purpose is to serve as an additional Cop to the target Cop.

            For every pieces of m-code which could not be mapped, translated or compiled, the COP calls the
            Matlab Engine.
            Needless to say, it does not exist today.


            Background                                                                                                                                  
            In 2005, when Catalytic Matlab Accelerator/Compiler could not tackle an exotic Matlab construct, it would make "a call into Matlab". The reason is simple, Catalytic was an add-on to Matlab so we could always use its resources as needed. The big drawback was very poor performances since Matlab interprets the code.
            Also since the last century The Mathworks offers a different type of Matlab engine called the MCR.
            Oddly enough, the concept is simple to understand. You develop a demo in Matlab. How to show it to customers who do not have Matlab? simple, you 'rebuild" your demo with the MCR and delivers an executable to be run on the customer PC. This is an over-simplification of the process, but the point is made. Matlab exists as a x86 executable (say < 1G) which can be ported, hence it could be hardcoded.
            Finally, this one is for me. Since i was a kid I made a point of honour at guessing April's fool jokes. But every 20 years or so I got really taken for a ride, especially if it is something "I knew they would be stupid enough to do that" Kind of Apple buying Tesla.
            And so this year, Clever Cleve came up with a Matlab announcement: "the-matlab-watch".
            http://blogs.mathworks.com/cleve/2015/04/01/experiencing-the-matlab-watch/?s_tid=Blog_Cleve_Category    Now, where he was really clever is that the platform was kind of realistic. He complained a bit about the lack of keyboard. Anyway he took me a few re-readings to realize it was a joke. Gosh, as John Updike said "an old cuckoo has always a second wind"

             The Near future: The Matlab Box                                            
            It is obvious that there is niche market for a Matlab box. The problem is "is it bigger than Wall street?". Still, I would love The Mathworks to come up with something even if it is only a re branding operation. At worst they would piss-off a few ecosystem customers. Instead we have a myriad of analytics vendors and solutions, which is fair enough but has no interest for me, at least in this Blog. 

            Sunday, May 20, 2012

            T.O.C 20 May 2012

            1. So you want to be a DSP architect?
              1. The long story
              2. Background Checks 
              3. Building Block issues
              4. Methodology and tools issues
            2. Arithmetic BB
            3. Bit wise BB
            4. Vector BB
            5. Math BB
            6. DSP BB
            7. Matlab BB
            8. Found on the web
            9. Stop me if you've heard this one before!
            10. APPENDIX - DSP Architecture 
              1. DSP of the First Kind (1980-95)
              2. DSP of the Second Kind (1995-2005)
              3. DSP of the Third Kind (2010- ?)
              4. DSP of the lost kind (1995-2005)
              5. DSP  of any kind (1950-2050)
            11. APPENDIX - DSP structural Building Block
              1. AGU 
              2. PCU
              3. DAU
              4. SU (The Shuffle Unit) 
              5. RF (Register File)
              6. Bit slice: Bit slice building blocks

            Bit Wise BB - an introduction

            Bit Wise BB
            This is  a heck of a topic!
            Firstly it overlaps ISA studies  and DSP Building Blocks in at least 2 specific places (the DAU and the shuffle units).
            Secondly, the variety of  function classes is very wide. For instance, pack/unpack, Count Leading Signs, Gallois Fields, bit manipulation.
            Thirdly, and this will serve as an introduction to this topic, the function themselves can become amazingly out of hand. 
            Finally, Matlab bit wise capability is limited or non coherent. As of this writing our latest bit wise library is floating point based!

            Background                                                                                                                                  
            In 1996, when we were designing the Tricore ISA, Bruce came up with a set of so called "permute" instructions and I do remember Rod's reaction as between bafflement and  irritation.
            On Bruce side, the truth is that all Media processors had this class of instructions [ ref: search Ruby Lee].
            On Rod side, we all know that the most precious resource in ISA design is the opcode. And, the biggest issue with bit wise instructions is that they are opcode hogs (always requiring dozens of control bits).
            For the sake of the story, I would add that, since 1996 I have been involved with this issue multiple times and seen a few people falling in this trap.     

            Example: TI C64+ : instruction PACKHL2                                                                                              
            To illustrate this topic, we will now discuss in details the PACKHL2 instruction of the TI C64+ DSP.  We will then expand to all C64+ PACK instructions.

            naming convention and syntax issue
            The instruction mnemonic PACK implies that the destination register is smaller than the source register(s), HL is a subtype and 2 in the TI terminology means 2 sub-word operations. Since the registers are 32-bit wide,  ADD2 (sub2, etc.. ) will mean 2 x 16-bit additions, and ADD4 will 4 x 8-bit additions in parallel.

            The syntax is
                  PACKHL2(src1,src2, dst)   where src1,src2, dst  are all 32 bit registers

             Question: since he syntax is exactly the same as  
                  ADD (src1,src2,dst)   where src1,src2, dst  are all 32 bit registers
            why do you need the suffix 2?
            Answer; In a 32-bit architecture, the 32-bit register is generic. All instructions use 32-bit registers. What matters are the operations inside the 32-bit registers. In this case 2 implies 16-bit data.
            Question:  PACK implies some kind of register demotion. The syntax dst = src1 <op> src2 is just like any standard 2-operator syntax. The 2 registers are operated upon and the result is written to destination. Where is demotion in this type of operation?
            Effectively, we would be more comfortable with   dst32= PACK(src64) where src64 is a 32-bit register pair (such as A3_A2). 
                                     PACKHL A3_A2,A0       ; a syntax using a 32-bit register pair (64bit )  
                                     PACKHL A2, A8, A0       ; TI syntax using  two 32-bit registers
            but the advantage of the TI syntax is obvious: register flexibility.

            Definition
            We are now entering the core of the matter. What is the definition of the PACKHL instruction? First we want something simple to describe this instruction. Writing a "C" definition is rather wordy and the TI standard description is more complicated than needed.  The simplest description is to consider the 2 registers side by side, src2 on the left (made of the two half words D_C) and src1 on the right (respectively B_A), the result will be C_B.  
                                  D_C B_A
                                 C_B       
            We have now the following definitions:
                                    Z = concat(hi(X),lo(Y));    % hi() and lo() are self explaining functions
                                    Z = [hi(X) lo(Y)];              % even more Matlab like
                                    Z = [X(31:16) Y(15:0)];     not so Matlab

            And they all look clear. But concatenation is not the same as packing. In fact, intuitively it is the contrary. One increase the variable size, the other reduce it.

            Towards a general definition of PACK
            Let us start with a most general definition. The register size is 64-bit and the granularity is 8-bit. Both values are very reasonable in 32-bit architectures.We will call this instruction PERM(ute).
                                                  PERM(src64, dst64, controlword);
            We will see later that PACKHL2 is just a sub-case of PERM.
            PERM is easily described. as shown the following example
              src64   HGFEDCBA
                     8x8 switch
              dst64   AAGHFEDC
            In this description each  letter represents a byte and note the little endian choice.
            But then what is the control word and how many bits do you need? The number of bits is the problem. In this case we have for destination a 64-bit register made of 8 sub-component (bytes). Each byte can receive any of the 8 source bytes (a 3-bit control) and since there are 8 destination bytes, the total number of bits is then 3x8= 24 bits. Not an easy decision to make in a 32-bit opcode! And using a register to hold the control word is a poor solution. It means a 3-cycle instruction (MVK, MVHK, PERM).
            Finally the control word syntax is very straightforward. We just use the representation of the destination register.In the example above it is:
                                          PERM(src64,dst64, AAGHFEDC)

            Matching PERM and PACK
            Using the similar definition to the one above, it can be seen that PACKHL2 is equivalent to:
                                                     PERM(src32,src32,dst32, FEDC)
             In fact we can match all C64+ PACK Instructions in the same way (see table) 
            It must be noted that the C64+  DPACK instructions are effectively more like PERM than PACK since the sources (combined) and destination have the same 64-bit width.


            Conclusions
            • Defining bit wise functions just by looking at the datapath is relatively easy and can give simple yet very powerful structures. Architects (attracted by elegance) will always love that.
            • The problem is the control path which can become rapidly out of hand.
            • To illustrate the problem we took an intruction from the C64+ instruction (PACKHL2) and extended it to a generic Permute instruction. The number of bits require to describe the instructions would be 24. 
            • Thiis problem applies the same to a building block or a coprocessor unit. 
            • With reference to the C64+ we will now compare the two approaches:
              • Advantages of using a general PERM instruction
                • Conceptually, it is very simple.
                • Software implementation is very direct.
                • Very flexible: any source byte can go to any destination byte. Any source byte can be duplicated (or replicated n times) in destination. 
                • No need to do long studies and have drastic selection to choose the right PACK datapath to implement (this is very often the case with bytes). 
                  • C64+ offers only 2 choice: byte even and byte odd
              •  Shortcomings for using a general PERM instruction
                • Need 24 bits of control. This is not realistic in 32-bit ISA.
                  • C64+ defines only 8 instructions.The footprint is minimal
                • Having 24 bits gives power(2,24) possibilities; how to test that? (by construction?)
                • The flexibility advantage may be a delusion. Some features are missing. For instance, just looking at the C64+ ISA (sign extension, packing with saturation). saturation.
            • While TI made the right choices for the C64+ ISA, a different situation (64-bit ISA, dedicated COP, etc..) might give different results.The astute reader, we are sure, has already plenty of ideas and solutions.
            • BUT this is not the point of this section. The point is to make sure that you understand the main risk associated with Bit wise instructions (the control bits) . To be forewarned is to be ... 






              


             



            Saturday, February 18, 2012

            SDR (Software Defined Radio)

            SDR (Software Defined Radio)
            This section covers an ambitious project. Let us imagine a wireless terminal made of 2 CPU blocks:  a powerful host armed with all media capabilities and a second block which can process any kind of RF signals and radio  protocols  (i.e 3G,4G, Wifi, gps, up to future standards such as cognitive radio). This second block is commonly called Software Defined Radio (SDR).

            Background
            The concept of SDR is as old as the radio [ref 1] but in our context we will put the burden on the solid shoulders of Les Mintzer circa 2000 [ref 2]. This was an FPGA implementation which makes a lot of sense as the money and impetus came from the military. In the same school see also Xilinx's Chris Dick, Spectrum's Lee Pucker and a few others [ref 3,4,5,6,7,8,9,10]. Now at this stage their implementation of SDR was far from fitting the definition but it made sense. Worse was to come. Firstly, in a brilliant case of taking the tree for the forest, a bunch of guys were using the PC/pentium as their example of SDR.They could do everything except radio.
            Frankly first time I heard about that, I thought of genius generating radio from a PC  by using the EMC radiation of the chips, this sounds silly but after all this was the time of UWB so why not?
            And then we had the usual silicon valley trend: a new buzzword for a new type of chip architecture [ref 11,12,13].


            Reference
            1. Dan Stransberg  " A century old technology enters the digital age", EDN , October 28 1999
            2. Les Mintzer "Soft Radios and modems and FPGA" CSD mag, february 2000 p.52
            3. Chris Dick, Fred Harris "the platform FPGA Software defined Radio" good paper but no reference, likely a conference circa 2001
            4. Chris Dick "FPGAs cranked for software radio" EET, April 10, 2000
            5. Chris Dick " A case for using FPGAs in SDR Phy" EET, August 12,2002
            6. Lee Pucker "Paving Paths to SDR" csd mag june 2001 p.19
            7. Lee Pucker, Systems Architect, Spectrum Signal Processing "Distributed Architecture for SDR" CSD 2001
            8. Robert Sgandura PM Pentek " W221: Software Radio - From concepts to Implementation" likely at Communication System Design conference 2000 or 2001
            9. Cord Finlay "Understanding SDR requirements" Wireless System Design July 2001
            10. "Soft Radio key to universal applications" EE Times Special Report March 19, 2001
              1. so full of hope, fortunately Loring Wirbel deflated the hot air in EET aug 12, 2002.
            11. John Ralston "SDR emerges to address wireless industry needs" Wireless System design october 2000 
              1. Describes the Morphics architecture
            12. Armin Nueckel, Prashant Rao "The use of Reconfigurable Processor Arrays in wireless Infrastructure Systems" July 25, 2001 {white paper?} 
              1. Describes the PACT architecture
            13. Paul Master "the baseband solution for the worldphone" www.quicksilver.com 2001

              Saturday, February 11, 2012

              DSP of the lost kind

              DSP of the lost kind
              The goal of this section is to laugh through the multiple attempts at defining new DSPs , next craze DSPs, further than DSPs, beyond DSPs , everything and their contrary as we went from 1995 to 2005 . Among the most notables we will have a quick pass at: Media processors.
              Now, the obvious question is: why should we care? Once more, the answer is that there is a very large difference in use model between a  GP full market CPU and a COP.
              What is ridiculous in one case becomes brilliant and  maybe most effective for the given application (see NVIDIA evolution). Also some of the brightest architects of  our time were involved in the design of these new machines. Finally this is still one of  the riches vein of ideas available publicly (but not freely) (IEEE, ACM, MDR, etc..). 

              BACKGROUND
              There are two reasons why we feel sad about this "Idiot Wind" period .
              1. Too many people were getting carried away and too many arguments just did not make sense. 
                1. For instance we computed it would take 1200 men years of application software (not tools)  just to meet the minimum requirements of fulfilling the claim of being a Media processor.
              2. For a while, we went in the same direction, (and we have to thank the influence of Jim Turley for that) (mind you he changed his mind pretty quickly too).
              DESCRIPTION AND TYPES
              • Type 1
                • Video Signal Processors
                • Video DSP
                • Media Processors
              • Type 2  Legoland
                • Multi IP DSP, Cluster Based DSP
                  • Bops, Cell, Improv, Equator
                  •  atsana, chipswright, clearspeeed, craddle, e-lite, powerFFT, sandbridge, siroyan, synputer, telairity, tops, trips  
                  • Infinite-tech RADarray  (neil stollon ICSPAT 99)
                    • RADarray of RADcore coprocessors connected to host
                    • Host and peripherals are 3rd-party IP
                    • Each RADcore is a reconfigurable stream algorithm coprocessor.
                    • Each RADcore is made of user selected Execution Units (EXUs) interconnected by reconfigurable data path bus architecture
                    • EXUs form independent elemental processing blocks such as ALUs, MACs, memories, I/O,etc..
              • Type 3 Domain DSP
                • Wireless DSP,
                • Speech, Audio DSP
                • Broadband DSP
              • Custom DSP, configurable DSP, 
              • Reconfigurable DSP
              Media Processor (MP) (or MVP) (V for Video)
              Born around 92-94, the  MVP had a peak around 1997, went through multiple "down and up"s before final death by 2005.The most famous (first wave) names were TriMedia, MicroUnity, Chromatics, NVidia [ref.1]
              They had several characteristic in common:
              • It was a new type of processor (the MEDIA processor)
                • from our DSP perspective it was a bit baffling. Why not use a DSP?
              • They were the latest incarnation of the MultiMedia (MM) craze.
                • very soon followed by MM extensions. Compare MP and MM extension
              • For some reasons most were based on VLIW.
                • Why?
              • While 5 of 7 MM functions could be done by a DSP,  they rightly believed that a DSP would not be able to do the last 2 (video and graphics). 
                • But then why do they think they could?
                • What were the 7 functions?
              • Finally, the most obvious characteristic was their disdain for the Pentium. In a typical Silicon Valley Fashion, after loosing the Risc war, that was the new  frontier.In a nutshell:
                • First they lost the PC seat to the Pentium (that was CISC versus RISC)
                • Then they (re lost) the PC MM seat to the Pentium (that was GP processor versus Media processor)
                • Finally they could not even find a small coprocessor seat in the PC, i.e. to be used as a Media COP!!
                  • This is the point which merits attention. See NVIDIA below
                • In their final days, to add insult to injury they were totally inadequate for the embedded market. 
                  • Mind you, good luck trying DVD in 1999 or cell phone in 2000 with a 288 bit wide opcode and no software support!
              NVIDIA
              In the most remarkable feat of the century, NVIDIA changed direction by focusing on graphics instead of a rainbow of applications [ref.2].
                     Not sure about their recent direction towards general purpose, people never learn it seems!
              TriMedia
              In another example of adaptation, the champion of VLIW, after years of trying software optimization instead redesigned their chips with two multiple powerful video coprocessors. Also important, Phillips concentrated on platforms and Trimedia became "the other core" of the Nexperia platform. Instead of ARM+DSP it was MIPS+TriMedia. Mind you, the latest instantiation of the platform could be ARM + COPs.
              Definitely a good DSP core, the TriMedia  was a relatively simple 5 issue machine, had some neat tricks (fusing of  operations at the register file ports) and very eearly code compression. [ref 6,7,8,9]
              MicroUnity MediaProcessor
              Original coiner of the term, famous for John Moussouris, MicroUnity was the first on the scene and the first to die; they set the architecture standard for a single core parallelism pretty high. [ref 10,11]. They had a following with Equator.
              Chromatics MPACT
              The other Media Processor, noticed for its 792-bit datapath (if today it sounds peanuts, at the time the standard was 64-bit)... currently dead. [ref 12]
              Offer from the East
              As can one expect, Japan and its all powerful integrated consumer industry were not the last to follow the hype with their own offering
              - Sharp DDMP (1997)
              - NEC MP98   (MPR March2000)
              - Toshiba MEP (after giving up MPACT) (EPF 2002)
              - Fujitsu (was their VLIW a Media Processor?)
              - Hitachi ??
              - Matsushita ??
              Among the Noise
              While not serious commercial products, the MVPs developed by Universities and Research labs had interesting features. These are 2 examples we were familiar at the time but there are litteraly 100s of them [ref 13,14].
              • University of Hannover HiPAR  {search Johannes Kneip} (IEEE Video for Circuit and systems 1996)
                • Impressive X,Y memory
              • Infineon VIP {search Uli Ramacher} (Hot Chips 13) 
              Finally the true MVP
              The TI C80 also known as MVP (1993) remains an excellent architecture example. It is NOT a Media processor. Its was designed like any other DSP, with an ambitious target in MIPS which happened to be  video conferencing. Now as part of the strategy obviously the target was anything in the same order of MIPS magnitude. Most interesting it was a host+4xDSP core solution. Even more: how can you be so wrong in the memory model? (refer to somewhere else in this Blog). 

              Lessons Learnt
              NVIDIA
              Headline1: MultiMedia processor becomes UniMedia processor.

              Headline2: Media processor becomes Graphics solution..
              And they had enormous success as Pentium "Coprocessor" .
              Now this is an important lesson, for co-processing.  As a general purpose co-processing MM chip, the MVP architecture was a complete failure. As a "very focused" co-processing solution point, it is part of the standard PC architecture. 
              Microunity, Chromatics, TriMedia
              It would be silly to forget these architectures. On one hand, more recent and better architectures have been developed but on the other hand they are not so visible. 
              Further: the GPU story
              Obviously these are still questions (we have to wait another 3-5 years to learn the lesson):
              - KEY: is there a place for another type of computing (the GPU)?
                     One has to bear in mind that in 40 years of  silicon computing  (except for a short time of DSP) all attempts to compete with the general purpose model ended up in abject failures.
              - what kind of the processing model is the dual head(CPU+GPU) adopted by AMD ?
              - in the same way, what kind CPU+GPU is the Apple model?
               
              References
              1. John.A.Watlington "Video signal processors" , circa 1997, http://wad.www.media.mit.edu/peole/wad/vsp/node1.htm from a web site which (as usual) has gone pining for the woods. Too bad , it contained an excellent and succint table of comparison.
              2. Section news" NVIDIA changes direction" EBN December 23,1996 issue:1038 
              3. Bernard Cole "New processors up multimedia's punch" EET February 3, 1997.
                8. Lee, W., Kim, Y., Gove, R.J., and Reed, C.J., “MediaStation 5000: Integrating Video an
                1. it is more the answer from CPUs (MM extensions) to MM processors.
              4. Jim Turley "Multimedia chips complicate choices" MDR Feb 12, 1996 page 14
              5. Maury Wright "Media Procesors target digital video Roles" EDN  sep 1, 1998
              6. Tom Halfill "Philips Trimedia goes Mobile " MPR dec 5, 2005 
              7. Gert Slavenburg and his biking pal "DSPCPU operations for TM1100" 1999, Appendix A, Preliminary information
              8. Gert Slavenburg, etc..'Custom operations for Multimedia" Chapter 4 of the same reference
              9. Peter Clarke " Compressed VLIW meets multimedia" EE Dec, 1995
              10. Craig Hansen "MicroUnity MediaProcessor Architecture" IEEE micro, 1996
                1. quote "A broadband mediaprocessor extends and streamlines a general-purpose computer system to attain the goal of communicating and processing digital video, audio, data and RF signals (SIC!) at broadband rates using compiled, downloadable, software rather than special-purpose hardware. The instruction set, system facilities, and initial implementations of an architectural family of broadband mediaprocessors are introduced, and compiled software development is illustrated with an example and description of the development environment".
              11. John Moussouris (easy to remember mouse+souris) "A roadmap of the Mediaprocessor Design space" Microdesign Resources Dinner May 9, 1996
              12. Yong Yao " Chromatics's Mpact2 boosts 3D, MPR Nov 18, 1996
              13. SSC 27 12, dec 92 p1886
              14. SSC 29 12, dec 94 p1474
              15. http://www.cse.fau.edu/~borko/Chapter21_mc.pdf
                1. Just added this morning, just wished I had  read it before.kind of complement to our MM splash. Hey Borko is it public? 
              Answer to questions
              The 7 functions:
              • graphics including 3D
              • video (compression, editing , conferencing)
              • high quality Audio
              • computer telephony, Speech
              • communications, Fax, Modems
              • real time 3D games
              • DVD?