jueves, 23 de mayo de 2024

A tribute to four programming languages

This post is a tribute to four programming languages: FORTRAN, PASCAL, C++, and LISP.

The algorithm that was programmed calculates the roots of the second degree polynomial:  

ax2+bx+c=0


The document containing the four codes: "programa raices.doc", was written in August, 2001, and it is available in GitHub .

Enjoy!


lunes, 8 de abril de 2024

Utilizing Ordinary Differential Calculus to Detect R-Waves in Standard EKGs

 


The electrocardiogram (ECG or EKG) is a vital tool for analyzing heart health. It captures the electrical activity of the heart, with distinct peaks and valleys corresponding to different stages of the heartbeat. 

Willem Einthoven invented the first practical electrocardiograph in 1895 and received the Nobel Prize in Physiology or Medicine in 1924 "for the discovery of the mechanism of the electrocardiogram".

The EKG signal typically consists of several waves: P-wave, QRS complex, and T-wave.

One crucial element of an EKG is the R-wave, which represents the peak of ventricular depolarization: the moment when the main pumping chambers of the heart contract. Detecting R-waves accurately is essential for interpreting cardiac rhythms and identifying abnormalities.

The R-wave typically occurs approximately 200 to 300 milliseconds after the onset of ventricular depolarization. Using the location of the local minima in the first derivative, we can estimate the time of arrival of the R-wave.

While there are numerous methods for R-wave detection, employing ordinary differential calculus can provide a robust analytical approach. 

Let's explore how ordinary differential calculus can help us pinpoint these R-waves.

DISCLAIMER

This approach was tested only in an academic environment, as part of a lecture about Electrocardiography. See Course below.

The Math Behind the Beat

An EKG recording can be thought as a function of time, with voltage on the y-axis and time on the x-axis. The R-wave corresponds to a peak in this voltage function. In Calculus terms, a peak signifies a maximum of the function.

The first derivative of the EKG function represents the rate of change of voltage over time. At the peak (R-wave), this rate of change will be zero. So, finding the maximums of the first derivative will lead us close to the R-waves.

EKG signals are often noisy. The first derivative might have minor fluctuations around the true maximum. To refine our detection, we can take the second derivative. The second derivative represents the rate of change of the rate of change (think acceleration).

At the R-wave, the voltage is at its maximum, so the first derivative is zero. But just before the peak, the voltage is rapidly increasing, resulting in a positive second derivative. Just after the peak, the voltage starts decreasing, leading to a negative second derivative.

Therefore, identifying points where the second derivative transitions from positive to negative will pinpoint the exact location of the R-wave with greater accuracy.

Putting it Together

By analyzing the EKG signal with these derivative properties in mind, we can identify R-waves:

  1. Find Local Maxima in First Derivative: Scan the first derivative for points where the value reaches a maximum. These points correspond to potential R-waves.
  2. Verify with Second Derivative Minimum: For each potential R-wave identified in step 1, check the corresponding point in the second derivative. A true R-wave will have a minimum value at that point in the second derivative.

Benefits and Limitations

This method offers a simple, calculus-based approach to R-wave detection. However, it's important to consider limitations:

  • Noise: Real EKG signals can be noisy due to muscle movement or electrical interference. These can introduce false maxima/minima in the derivatives, requiring additional filtering or noise reduction techniques.
  • ECG Variations: EKG morphology can vary between individuals, and some abnormal heart rhythms might not exhibit the classic R-wave characteristics. More sophisticated algorithms might be needed for robust R-wave detection in such cases.

Conclusion

Understanding R-waves is essential for EKG analysis. By utilizing the concepts of maxima and minima in ordinary differential calculus, we gain valuable insight into the electrical activity of the heart and its rhythm.

Ordinary differential calculus provides a valuable tool for understanding and analyzing EKG signals. By focusing on maxima in the first derivative and minima in the second derivative, we can pinpoint the crucial R-waves, offering valuable insights into heart function. 

However, it's crucial to acknowledge the limitations of this approach and consider real-world complexities for robust R-wave detection in medical applications.

Then, further research and validation are necessary to refine and optimize this method for real world applications.

You can find two versions of the source code: 1) MatLab/Octave, and 2) Python in my GitHub repo Signal/EKG/R-Wave 

Course lecture: 

Cursos/Electro_Medicina/elecmed05_2 Electrocardiografía.ppt

Note: This presentation is written in Spanish.

Recommended lectures

1.  Bronzino,J.D. (Editor)  “The Biomedical Engineering Handbook, 2nd Ed. IEEE Press, 2000 Chapter 13 “Principles of Electrocardiography”

2.  Carr,J.J y Brown,J.M. “Introduction to Biomedical Equipment Technology” Chapter 8 “Electrocardiography” pp 197-233

3.  Del Aguila, C. “Electromedicina” Ed. Hasa, 1994 Capítulo 8 “Bases de la Electrocardiografía” (pp 129-159) y Apéndice III Electrocardiógrafo (pp 475-499)

4. Webster, J.G. (Editor) “BioInstrumentation”, 2003



martes, 27 de marzo de 2018

Why Data Scientists prefer Python?

I usually write code in Java and C. Then, to start answering this question, I performed a simple test by writing the same code in the three programming languages: C, Java, and Python. My goal was to understand the differences among performance of of the three approaches, specifically regarding CPU and memory usage.

Then, I selected an algorithm that can be easily ported to Python too: Retrieving the second maximum value among three random positive integers. In order to get valid timing measures, this process was repeated 1000000 times. (+)

It must be pointed out that this is not a real benchmark test, because that wasn't what I was looking for. Even, the code does not make any input/output or network operations.

All tests were done in the same environment(!):
$ uname -a
Linux <this> 4.15.10-300.fc27.x86_64 #1 SMP Thu Mar 15 17:13:04 UTC 2018 x86_64 x86_64 x86_64 GNU/Linux 


The executables were created by using the following tools:
C:
$ gcc --version
gcc (GCC) 7.3.1 20180303 (Red Hat 7.3.1-5)

Java:
$ javac -version
javac 1.8.0_161
 
$ java -version
openjdk version "1.8.0_161"
OpenJDK Runtime Environment (build 1.8.0_161-b14)
OpenJDK 64-Bit Server VM (build 25.161-b14, mixed mode)

Python:
$ python --version
Python 2.7.14


The size of the codes, excluding JAVA Virtual Machine and Python Interpreter, are:
C: 8568 bytes, about 35 lines of code.
Java: 2721 bytes, about 40 lines of code.
Python: 1486 bytes, about 30 lines of code.

The commands to run the programs were:
C:
$ ./secondmax_c 1000000
Java:
$ java -jar SecondMax.jar 1000000
Python:
$ python secondmax_p.py 1000000

The results, in seconds, are:
C:
Best time: 0,038512 s
Worst time: 0,041267 s
Average time of 10 samples: 0,039393 s (*)
Java:
Best time: 0,190164966 s
Worst time: 0,213463167 s
Average time of 10 samples: 0,1982906219 s (*)
Python:
Best time: 14,326186 s
Worst time: 14,948839 s
Average time of 10 samples: 14,6719912 s (*)


So, if these tests show that the worst performance in C was around 347 times faster than the best performance in Python, and the worst performance in Java was around 67 times faster than the best performance in Python: Why do data scientists prefer Python? 

The answer is straightforward, and it might sound pretty obvious, even before testing: Python is easiest.

C and Java demand professional programming skills that must be acquired after a personal process of growth that sometimes may be slow.

However, a data scientist with her/his computer can understand and develop an acceptable Python code quickest.

Furthermore, Python has many libraries and solutions that can be freely included in the applications, reducing the development time, and therefore, the man hours getting a valuable answer to the different issues.

Anyway, I still have two open questions:
1.- What should be the best performance's option from System Administrator's point of view?
2.- Which one is the choice if the target environment of the application is an embedded system? In example: 454 MHz ARM 9 CPU with only 128 MB of RAM?

Notes:
(+) If you want testing my codes in your environment, just write a comment below and I'll send these. 
(!) To avoid comparing apples to oranges, I have omitted in this post the results after testing the three codes in different environments. It was tested too the performance of the C version in an embedded system: 454 MHz ARM 9 CPU with only 128 MB of RAM; and the performance of the PySpark version in a Cloudera 5.12.0-0 cluster.
(*) The average performance was only included because this is a standard measurement. I think that the value of the "worst time" is the most accurate predictor of the real time response capabilities of an executable program.


 

lunes, 30 de octubre de 2017

The Sentient Enterprise

I just have read the innovative book "The Sentient Enterprise", written by Oliver Ratzesberger and Mohanbir Sawhney.

This book is a "must read" for all of us involved in the improvement of the decision processes.

To clarify why, it is worthy to mention only two fundamental IT problems that are analyzed inside the book:
  • Why the enterprises are wasting human resources, time, efforts, their computing power, and money by unnecessary duplicating data?
  • Why the IT analytic solutions are usually delayed, and many times become obsolete when they are finally finished?
In order to help solving these issues, Ratzerberger and Sawhney introduce a model with five stages:
  1. The Agile Data Platform
  2. The Behavioral Data Platform
  3. The Collaborative Ideation Platform
  4. The Analytical Application Platform
  5. The Autonomous Decisioning Platform
This approach decomposes the huge problem: better decisions, and allows to develop specific solutions for each stage; thus making the whole process affordable.

I specifically recommend reading in detail chapter 7 "The Autonomous Decisioning Platform", and chapter 8 "Implementing Your Course to Sentience".

These chapters expose important guidelines, and introduce a perspective that matches most of the research topics behind the Autonomous Intelligent Systems (AIS)

lunes, 25 de septiembre de 2017

How useful are Recommendation Systems?

Typically, recommendation engines and systems enhance the user experience, because they assist us in finding information, reduce search and navigation time, and increase our satisfaction. 

However, I am still receiving recommendations about options to buy vacation packs, books, movies, music, etc.; that I had reviewed more than one year ago. Even worse, I receive friend suggestions because they are friends of someone that I know… Why? Really, most of the time I am not interested in receiving those kinds of recommendations… 

Therefore, I asked myself: How many people accept and follow these recommendations? 

Unfortunately, I don’t have access to all the required data in order to evaluate this; but according to Symeonidis and Zioupos (“Matrix and Tensor Factorization Techniques for Recommender Systems” ISBN 978-3-319-41356-3):
  • Amazon – 35% of product sales come from recommendations in Amazon.com  
  • Netflix – 66% of movies rented in Netflix.com are recommended  
  • Google – 38% more click-throughs are generated from recommendations in Google news
In my opinion, the main reason of this behavior can be traced back to the origin of the algorithms that have been used to create recommendations: clustering, ranking, scoring, pattern matching, etc. This means “the History”, usually understood as Big Data, OLAP, OLTP, etc. But it is clear that this "historical approach" demands many resources: storage, computing power, and time. 

But: What happen if we’re looking for recommendations about the outcome of a “random” process? 

The situation becomes harder if we don’t have enough information about the process itself. Let’s put it on an easy way: We need recommendations that could not be related to the previous history of the process. 

Example: A recommendation with probability 0.7 is not a winning one in gambling. Unfortunately, the number of alternatives experiences an exponential growth in order to achieve a greater probability. It is difficult to create such recommendations using "brute force" algorithms, and the task will demand the most powerful computers. Even worse: There is not heuristics to "prune" the decision tree.

This means that the inputs to the recommendation’s process are not well suited because the data could be either too poor or too much, the representation of the recommendation’s knowledge has not been identified as it should be, the inference rules could not be useful because there is not a previous experience about the behavior of the process, and the explanation about how the recommendations were created does not allow to identify the whole reasoning process: the domain is not well formalized. 

Then, I'm asking myself: How to predict the future behavior of a system whose previous history might not be relevant to the predictive process?

I believe that a new approach is needed to the recommendation’s process. This approach should redefine how to analyze the inputs, how to build new models of knowledge’s representation, how to propose different inference mechanisms that might not be well suited and computed by Turing’s Machines, many alternatives of solution, and explanations about how these recommendations were reasoned that may not match the common sense.

As a conclusion, I think that research in Artificial Intelligence should include more than machine learning, deep learning, and neural networks; because the focus of the problem should consider not only the data -the history-, but also the mechanisms to extract real knowledge of them: models, representations, inference rules, and explanations.

This path will lead us to “skilled intelligent systems”: solutions that can be transferred from one domain to other domains, and recommendations that will be really useful.   

viernes, 15 de septiembre de 2017

Comentario - Comment

Estimados lectores:

Cuando oficialmente dejé de trabajar como Profesor Titular en el año 2010, publiqué mis clases en Español para dominio público y sin fines de lucro en OneDrive < https://1drv.ms/f/s!ApPTAVJ07A-CgRQ2NbUjw6U1nDnB > .

A lo largo de varios años he visto dichos materiales referenciados y reproducidos por diversos sitios como SlideShare también sin fines de lucro. (Ejemplo en SlideShare: < https://www.slideshare.net/fuvylvp/almacenes-de-datos-olap-y-minera-de-datos?qid=2634f8ee-fb1d-4c57-aec0-a0add8ce77e6&v=&b=&from_search=1 > )

Sin embargo, hoy encontré un sitio que se declara "sin fines de lucro", pero que pide "apoyo al compartir y/o descargar" uno de mis documentos: "Aspectos Avanzados de la Tecnología de Objetos".

Personalmente considero que esta información debe ser conocida por todos.

Agradezco de antemano cualquier comentario ó sugerencia sobre acciones futuras al respecto.

Dear readers:

When I officially quit working as a Professor in 2010, I published my lectures in Spanish for public domain and non-profit use in OneDrive <https://1drv.ms/f/s!ApPTAVJ07A-CgRQ2NbUjw6U1nDnB >.

Over the years I have seen such materials referenced and reproduced by various non-profit sites such as SlideShare. (Example in SlideShare: <https://www.slideshare.net/fuvylvp/almacenes-de-datos-olap-y-minera-de-datos?qid=2634f8ee-fb1d-4c57-aec0-a0add8ce77e6&v=&b=&from_search= 1>)

However, today I found a site that is declared "non-profit", but "needs support to share and download" one of my documents: "Aspectos Avanzados de la Tecnología de Objetos".

Personally I consider that this information should be known by everyone.

I appreciate in advance any comments or suggestions on future actions in this regard.

Dr. Juan Jose Aranda Aboy


lunes, 12 de junio de 2017

Building an intelligent Chef

Let's start by analyzing the scope of the job: Building an intelligent machine means that the system must pass Turing's Test. Therefore, we are introducing by default a "common sense" rule in our research.

An interesting problem appears if we want to create an intelligent machine that can act as a Chef.

Cooking is an open problem, and it introduces some challenges to the intelligent machine. We will consider only three:
1. Should it cook by replicating a recipe step by step? What about measures? Cooking time? Is Fuzzy Logic the tool that could help solving these problems?
2. Could it create its own recipes by using the available products only? Could it transfer some skills that it has learned before such as music composing or poetry? How? Would Evolutionary Algorithms help?
3. What results can be accepted as "tasty meals"?


We can continue writing challenges that our "Chef" should overcome, but the last one is the main problem: the expectations about the resulting meal vary for each person, and even worse: each culture redefines cooking according to its history, location, and standards.

Therefore, which one is the appropriate "output"? I'm afraid that there are many solutions, and all these "tasty meals" would pass Turing's Test, but some of them would not pass the people’s taste.

The next step should be to "model" our Chef. To do so, we will analyze the problem again by reviewing the required actions to build an intelligent machine transforming the inputs (i.e.: beef, salt, lemon, onion, garlic, margarine, etc.) in an output: the meal (Steak!).

First, the desired recipe should be selected. This can be done by surfing the Internet. Therefore, our Chef can solve this easily.

Second, it should be verified if there are all the required products in the recipe. Also, the intelligent machine must check if there is the required quantity of all those products. The cooperation of another intelligent system is needed to keep records of the existing products.

Third, preparation: the beef must be sprinkled each side with salt. Then it must be added lemon juice, garlic, and onion. So, the intelligent Chef will need some additional devices:
- Sprinkler,
- Squeezer, to extract the lemon juice
- Peel the garlic clove, and chop it
- Cut the onion 
Fortunately, we can assume that have been created a set of intelligent machines that can solve these problems.

Finally, the intelligent Chef should melt margarine in a large skillet over "medium-high heat", fry the steak on each side, and transfer to a hot serving plate. Two comments: 
A) Which one is the appropriate temperature of the skillet? What is the meaning of medium-high heat? 
B) How much time is needed to fry the steak on each side?


There are many other problems that the intelligent machine may ask itself. A few examples are:  
- What if there is "not enough" (less than the required) margarine, but there is enough vegetable oil and butter?
- What if the recipe must be cooked without garlic and onions, or less salt because of the consumer's requirements? 
- Can be used the same procedure to cook fry chicken?

This is only a preliminary exercise. Thus, I have written more questions than answers. However, I feel confident about the future, because Artificial Intelligence is still young, and there are many researchers contributing to the field, so my expectations are high.