Showing posts with label data. Show all posts
Showing posts with label data. Show all posts

Thursday, March 18, 2021

Structures: Scaffolds for growth

 

     For many startups, the total rule is "lean, green, mean". Do what you most need to do, as fast as you can, with as little excess, or non-mandatory, work as possible. When I was co-founder of our company, it was not unusual to be working 80-hour weeks. We knew what we had to produce and we had a few methods to try to get it into the hands of people who would pay us for them. (We soon needed to expand that "marketing and sales" aspect of the business.)

     That is the basics of trade in a nutshell -- produce what is of value to others who will give back things of value to you.

     That works in a barter economy as well as in an industrialized, capitalistic, economy. It is also true within other viable economic systems. As the business, or economy, or government, grows it can often end up "abstracted" where it is difficult to say exactly what things of value are being exchanged.

     Leap ahead and start imagining a business that has thousands of people working to provide tens of thousands of things of value and having to keep track of a hundred thousand purchases and transactions. If it all works smoothly then it could be done in the same manner as when the trade was just between you and someone else -- that simple, basic, barter agreement.

     But this is reality. There are few one-to-one relationships between any person and any proceeding from start to finish. Person A does one part to process C and Person B does something different to process G which directly influences process C but A has no direct visibility to process G.

     Confusing? Absolutely. And this is still only a very simple situation. There needs to be some type of documentation -- method of communication -- between Person B and Person A that provides insight into relevant aspects of Process G without inundating Person A with all of the other knowledge and systems that Person B is handling.

     So, simple transactions have need of simple processes. High numbers of interrelated transactions, people, and processes have need of much better access to, and keeping track of, relevant information. How do you succeed in growing the business from simple to complex?

     The base answer is "structures" which can be loosely defined as ways to organize information about what is being done, (who/what/when/where/why) by the people who originate the information, to have it accessible to those who need to know that information. The other leg is "processes" -- which is involved with how that information is processed, saved, distributed, and otherwise not lost in the cascading effects of a successful large business.

     Processes can (and do) make use of various apps and programs. But without structures, the processes cannot do much of anything because they don't have the data with which to use those processes. Also, processes differ with every aspect of the business. A process for generating ideas. A process for estimating, and keeping track of, work. A process for manufacturing inventory and supply control chains. And so forth. But all of the processes rely on structures.

     The primary importance of determining just what information is needed for the business is that it remains approximately the same no matter how large the company gets. (Yes, as a business reaches certain growth points, new regulations may come into play.) This facilitates growth. The information has to be there but, when the company is small, it can be retained within various people's memories. Just like it seems to be true to a teenager, all employees of a startup are deemed to be immortal and those valuable data are always available.

     Absurd? Certainly. But it is so very easy to eliminate those items, that seem to not immediately affect the bottom line, when you are small, focused, and overworked. Resist. The data can be written on a large notepad or, for transitory data, on a white board. But get it written down. As the company grows, you are going to run out of room on those notepads or they will become too many to search through easily. So, you develop (or obtain) new processes and applications that help you to manage that data. But you already are used to getting, and documenting, that data. You are prepared for growth.

     

Thursday, July 11, 2013

BIG data and data mining

In my household, big data is most directly related to the piles of LEGOs (or LEGO-system building components) that my boys have scattered around the house. Needles in haystacks are more often used as examples. My library of books around the house would be another example. In each case, big data basically means a lot of data.

A lot of anything, of course, is subjective. There are thousands of pieces of straw in a haystack. There are a few thousand books around my house. My boys have a couple of thousands of LEGOs. However, in the world of business (and surveillance) big data usually refers to hundreds of thousands (or even millions) of records -- each of which may have many minutes (audio) or many members (items sold in purchase records or words in emails, for example). Big data is just a way to describe lots of data.

Data mining is the process of finding that special yellow 2 by 2 LEGO in the pile, or finding the needle in the haystack, or finding a specific audio record that talks about things that are considered suspicious or dangerous.

Data mining has three basic components -- collection, storage, and analysis. These are not necessarily discrete stages but we'll discuss them separately (calling out exceptions).

As evidenced by the physical examples at the beginning of this blog, big data has always existed. Consider the stacks of paper birth certificates, or other historical documents that exist and which may need, from time to time, to be searched. The ability to effectively handle, and use, big data has gotten much easier since electronic formats have become standard.

  • Collection. Collection usually occurs at the time of transmission (when the originated data is moved to a destination). This might be a phone call. It could be at a point-of-sale (POS) cash register after the order has been finalized. It might be the registration record for a class. Collection may either occur at the intended destination (the company invoice/purchase order database) or via interception. Interception is where collection occurs somewhere other than the intended destination -- "wire tapping", people looking over your shoulder when you enter your credit card security information, and so forth.

    Collection can occur anonymously or personalized. Personalization basically means that the record is associated with a corporate or living entity. In the case of a sale at a grocery store, the data will be associated with that store (and, possibly, that cashier and cash register). If you use a credit/debit card or a store "club" card, then the data can (and probably will) be associated with the person in addition. Generally, anonymous collection is considered innocuous while personalized collection is not. This does not mean there are not "legitimate" (proper, honorable) reasons to collect personal data but it does mean that the person may have concerns as to the purpose and safety of the data.

  • Storage. This always occurs at some point. However, it may be transitory if the data are removed upon receipt and analysis. Consider a "normal" phone call. The audio message exists (and is stored) from the origination (talking) until the receiving person analyzes it. If the message is redirected (to voice mail, for example), intercepted, or copied, this may turn into a permanent record requiring long-term storage.

    Transactions (purchases, registration, email correspondence) where the data needs to be used in the future are almost all "permanently" stored. Of course, they can still be deleted in the future -- but, without advance knowledge of when, or if, this will occur they must be considered permanent.

  • Analysis. This can occur during the process of collection or it may occur later (after storage). Anonymous data is often analyzed statistically. How many of product X were sold by store Y in city Z? How many of product X were sold in state B? How long is the average voice call within a state? Trends can be analyzed over time. Store Y in city X sold NN of product X at price B. They sold GG of product X at price C (can be used to determine overall profit using margin versus quantity sold). Product F sells very well during the time period D through G but not very well in period H through M (seasonal item to be stocked differently depending on time of year).

    Analysis can also be personalized. Customer ABC buys a lot of product F. Product G is similar but there is a greater profit margin on G -- send Customer ABC coupons for product G to get them to start buying product G on a regular basis. Or Customer DEF only buys product F if the price is below $ZZ.ZZ. Customer BEF is now buying baby products -- notify baby supply companies of contact information.

    Finally, analysis can be triggered. Surveillance can use trigger words, or sequences of words (either written or audio) to divert records to further analysis. If you start buying diabetic-related foods and medicines, the data CAN be forwarded to your insurance company (and yes -- if the data is associated with you, then they CAN find your insurance company).
Big data does not change the stages but it does change the methods. There will often be multiple layers of analysis so that each step reduces the number of records to be analyzed. Analysis upon collection will specifically affect the manner in which the data are sorted and stored. And so forth.

People usually don't object to anonymous statistical analysis. They may start feeling threatened with personalized statistical analysis although they may also benefit from the results.

They often will feel threatened with triggered analysis because their "private" data are being used without explicit permission and can be used to exploit the data in some way. In addition, triggers can lead to false conclusions quite easily (you were actually buying diabetes supplies for your great Aunt, you have been reading a book about bad thing XXX and were discussing it with a friend). Big data methods are particularly susceptible to false initial triggers (although, hopefully, further analysis will filter more appropriately).

Friday, August 28, 2009

Computer Literacy 101 -- what are programs?


Data falls into two categories, as we saw in the previous blog. These categories are instruction data and program data. Instruction data can also be called a program -- which makes use of the program data to fulfill its purpose. Many times, a program will be called an application. Most people refer to programs as applications if they are widely used by different people. A word processing program may be considered an application. A spreadsheet program may be considered an application. Programs that people do not directly make use of are usually not called applications.

Programs consist of a sequence of instructions that tell the computer what to do. In the first blog of this series, we mentioned the types of instructions that may be given a microprocessor. These instructions are called machine language because they are sets of datum that are interpreted directly by the microprocessor. For example, the decimal value of 026002 will tell an HP 2100 computer to "jump" (change the current address for the next address) to address 002. The 026 portion is a "change the current instruction address to" code for the microprocessor and the 002 portion tells it the new value for the instruction address. Frankly, you don't really want to know much more than that unless you are involved with microprocessor design. Most microprocessors have their instructions written in binary (0s and 1s) or hexadecimal (symbols 1 through F for each numerical location) but the computers of the "early days" were not always standardized.

There is a hierarchy of languages used to program computers. At the base is machine language -- normally represented by a series of 0s and 1s -- like 10011001100011100011111101011110 -- which would be considered 32-bits. Current-day programmers almost never use machine language directly. The next level is called assembly language which uses a set of readable codes that can be directly translated into machine language. An example of assembly language might be "JMP START", where JMP is the operation to be performed and START is a symbol for an address that is the value to be used by the operation. The next level consists of many different languages in a family called high-level computer languages. A compiler changes the high-level language into assembly language (or, sometimes, directly into machine language. An assembler changes the assembly language into machine language. Finally (at least, at this current point of time) there are machine-independent languages which are used to create programs that may be run on many different computers without being changed.

Programmers write programs. Very few write machine language programs. More, but not many, write assembly language programs. Most write programs in high-level languages. An increasing number write programs in machine-independent languages (such as Java). However, all of them end up actually creating machine language -- with special programs such as Java interpreters/compilers, compilers, and assemblers acting to make it into this special, final, form.

I said earlier that some programs are visibly used by people -- and these are called applications. The ones that are NOT visibly used by people are sometimes called system programs. These are programs that enable to computer to perform the acts that people want. A printer program will be used to allow people to print a document from their application. At the core of all of the system programs is a particular program called an operating system and this will be addressed in the next blog.

Thursday, August 27, 2009

Computer Literacy 101 -- what are data?


All computers work with data. But, what are data? I say "what are data" because the word data is a plural one -- it is the plural of datum. However, almost no one ever uses the word datum and just about everyone treats data, grammatically, as singular. A datum is a single piece of information -- yes or no, it is raining or it is not raining, you have eaten breakfast or you have not eaten breakfast. You will note that a datum only indicates a yes/no or off/on, binary, condition. Most of the time, when we need information, it is really a collection of datum -- or data.

The same is true with computers -- they work on each individual datum but they pull them out of a pool of data. This data (I will use the conventional singular grammar here) is kept in storage, as I pointed out in a previous blog. It is then transferred from one storage area to the local RAM where the microprocessor can directly work with it.

Data is used by the microprocessor at many stages. The first stage, or startup (or bootup), is when the microprocessor first receives electricity. The actual hardware (the collection of semiconductor chips, and other discrete electronic components) is designed to start transferring data from a specific memory storage area and address. Often, this is address zero (0). This means that the microprocessor will transfer data from address 0 (the actual physical location, once again, depends on design of the hardware) to its working memory. It then executes the data -- it starts to perform specific operations based on the contents of that data that was at that address and then increments the address for the next instruction (usually by one -- unless the first instruction says something else) and then executes the operations for that address and on and on.

The next stage occurs once the registers (we talked about them in the storage blog) of the microprocessor have been filled with working data. At this stage, it is prepared to continue to execute instructions as it transfers them from memory. You can look at is as having two stages -- the first where the microprocessor "wakes up" with no specific contents in its registers and the second when the microprocessor has been initialized and can now proceed to work as the data tells it. Or, you can look at it as having three stages -- the first one the boot stage, the second is the startup stage where it is still getting ALL of the hardware connected to the computer ready to be used, and the third stage where any type of instructions can be executed because all of the hardware has been set up to be ready for use.

All of this works with the data. The data, for a general purpose microprocessor, is what makes it able to work differently each time it is turned on. For a specialized microprocessor, the incoming data starts the activities for which the microprocessor is designed.

Somehow, I managed to avoid the word program in this description but data are often split into two categories. Instruction data, or programs, are executable -- they contain instructions for the microprocessor while program data is used by the programs to produce more data. The difference between these categories of data is that instruction data does not change (unless someone specifically writes another program to change that instruction data -- the topic of viruses and software patches).

I'm going to swap the next two items on my original "computer literacy" list and talk about programs more in the next blog.

Smoke Gets in Your Lungs (updated)

     This is an article that I published in here on February 22, 2013. I try to make my articles “timeless” as I try to work with “foundatio...