viernes, 24 de mayo de 2024

A tale of programming languages Clojure

 A tale of programming languages: Clojure

It was tested ChatGPT Code Copilot to generate the function "solve-quadratic" from the picture Raices.PNG :



FINAL CODE

===========

(ns quadratic-equation

  (:require [clojure.math.numeric-tower :as math]

            [clojure.test :refer :all]))


(defn complex-number

  "Creates a complex number as a map with real and imaginary parts."

  [real imag]

  {:pre [(number? real) (number? imag)]}

  {:real real :imag imag})


(defn solve-quadratic

  "Solves the quadratic equation ax^2 + bx + c = 0.

  Returns a map with the type of roots and the roots themselves.

  

  Parameters:

  - a: coefficient of x^2

  - b: coefficient of x

  - c: constant term

  

  Returns:

  - A map with :type key indicating the nature of roots (:linear, :real, or :complex)

    and :roots key containing the roots.

  - If inputs are invalid, returns an error message as a string."

  [a b c]

  {:pre [(number? a) (number? b) (number? c)]}

  (let [compute-linear-root (fn [b c]

                              {:type :linear :root (/ (- c) b)})

        compute-real-roots (fn [b discriminant denom]

                             (let [sqrt-d (math/sqrt discriminant)]

                               [(+ (- b) sqrt-d) (- (- b) sqrt-d)]))

        compute-complex-roots (fn [b discriminant denom]

                                (let [sqrt-d (math/sqrt (- discriminant))

                                      real-part (/ (- b) denom)

                                      imag-part (/ sqrt-d denom)]

                                  [(complex-number real-part imag-part)

                                   (complex-number real-part (- imag-part))]))]

    (cond

      (and (= a 0) (= b 0) (= c 0)) "Infinite solutions"

      (and (= a 0) (= b 0)) "No solution"

      (= a 0) (compute-linear-root b c)

      :else

      (let [discriminant (- (* b b) (* 4 a c))

            denom (* 2 a)]

        (cond

          (pos? discriminant)

          {:type :real :roots (map #(/ % denom) (compute-real-roots b discriminant denom))}

          (zero? discriminant)

          {:type :real :roots [(/ (- b) denom)]}

          :else

          {:type :complex :roots (compute-complex-roots b discriminant denom)})))))


;; Unit Tests

(deftest test-solve-quadratic

  (testing "Linear equation"

    (is (= {:type :linear :root 2.0} (solve-quadratic 0 2 -4))))

  

  (testing "Infinite solutions"

    (is (= "Infinite solutions" (solve-quadratic 0 0 0))))

  

  (testing "No solution"

    (is (= "No solution" (solve-quadratic 0 0 2))))

  

  (testing "Quadratic equation with two real roots"

    (is (= {:type :real :roots [2.0 1.0]} (solve-quadratic 1 -3 2))))

  

  (testing "Quadratic equation with one real root"

    (is (= {:type :real :roots [1.0]} (solve-quadratic 1 -2 1))))

  

  (testing "Quadratic equation with complex roots"

    (let [result (solve-quadratic 1 1 1)

          roots (:roots result)]

      (is (= :complex (:type result)))

      (is (= {:real -0.5 :imag 0.8660254037844386} (first roots)))

      (is (= {:real -0.5 :imag -0.8660254037844386} (second roots)))))

  

  (testing "Extreme coefficients - large values"

    (let [result (solve-quadratic 1e10 -3e10 2e10)

          roots (:roots result)]

      (is (= :real (:type result)))

      (is (every? number? roots))))

  

  (testing "Extreme coefficients - small values"

    (let [result (solve-quadratic 1e-10 -3e-10 2e-10)

          roots (:roots result)]

      (is (= :real (:type result)))

      (is (every? number? roots))))

  

  (testing "Invalid input handling"

    (is (thrown? AssertionError (solve-quadratic "a" "b" "c")))

    (is (thrown? AssertionError (solve-quadratic 1 "b" 2)))

    (is (thrown? AssertionError (solve-quadratic 1 2 "c")))))


;; Run the tests

(run-tests)

===========

Initial prompt:

You are an expert programming Clojure. Can you write the best code in this language for the algorithm in the attached picture?

To implement the algorithm depicted in the image in Clojure, we need to solve quadratic equations of the form 𝑎𝑥2+𝑏𝑥+𝑐=0. The image shows the steps and different cases based on the discriminant 𝑏2−4𝑎𝑐.

Let's break down the problem:

  1. Linear case: If 𝑎=0, the equation reduces to 𝑏𝑥+𝑐=0.
  2. Discriminant calculation: For 𝑎≠0, calculate the discriminant Δ=𝑏2−4𝑎𝑐.
  3. Discriminant cases:
    • If Δ>0, there are two real roots.
    • If Δ=0, there is one real root.
    • If Δ<0, there are two complex roots.

Pseudocode

  1. Check if 𝑎=0:
    • If true, solve the linear equation 𝑏𝑥+𝑐=0.
  2. Calculate the discriminant Δ=𝑏2−4𝑎𝑐.
  3. Evaluate the discriminant:
    • If Δ>0, compute the two real roots.
    • If Δ=0, compute the single real root.
    • If Δ<0, compute the two complex roots.

Explanation

  • The solve-quadratic function handles different cases based on the value of a and the discriminant.
  • The complex-number function creates a map representing a complex number.
  • The example usages demonstrate solving quadratic equations with different coefficients.

Second prompt, following "Next steps" suggested:

Write unit tests for the solve-quadratic function to ensure it handles all edge cases. Add type hints and documentation to the functions for better readability and maintainability.

Explanation

  1. Type Hints and Documentation:

    • Added :pre conditions to ensure the inputs are numbers.
    • Documented the complex-number and solve-quadratic functions with parameter and return descriptions.
  2. Unit Tests:

    • Created the test-solve-quadratic test suite using the clojure.test framework.
    • Added tests for linear equations, infinite solutions, no solutions, quadratic equations with real roots, and quadratic equations with complex roots.
    • For complex roots, the test verifies both the real and imaginary parts of the roots.
  3. Running Tests:

    • The run-tests function runs all the tests in the namespace.

Third prompt, following "Next steps" suggested:

Refactor the solve-quadratic function to handle more edge cases or optimize performance. Add more detailed error handling and user-friendly messages in case of invalid inputs.

Explanation

  1. Refactor for Edge Cases and Performance:

    • Improved the condition checks for different cases.
    • Simplified linear equation handling by checking specific cases directly.
  2. Error Handling:

    • Added assertions to ensure the inputs are numbers.
    • Provided clear error messages for different invalid input scenarios.
  3. Unit Tests:

    • Expanded the test suite to cover additional edge cases and invalid inputs.
    • Included tests for invalid inputs to ensure the function throws appropriate errors.

Fourth prompt, following "Next steps" suggested:

Optimize further by precomputing values or using more efficient mathematical functions if necessary. Add more complex unit tests, such as handling extremely large or small coefficients, to ensure robustness.

Explanation

  1. Optimization:

    • Extracted common computations into helper functions (compute-linear-root, compute-real-roots, and compute-complex-roots).
    • Precomputed denom and reused it in the calculations.
    • Reduced repetition by defining functions for real and complex root computations.
  2. Error Handling:

    • Maintained assertions to ensure inputs are valid numbers.
    • Error messages are included for invalid inputs, ensuring user-friendly responses.
  3. Unit Tests:

    • Added tests for extreme coefficients (large and small values) to ensure robustness.
    • Included existing tests to cover different edge cases.
    • Invalid input tests ensure the function throws appropriate errors.

Next Steps

a. Further optimize by using Clojure's transients for performance if necessary. b. Extend the function to handle more complex scenarios, such as higher degree polynomial equations.


jueves, 23 de mayo de 2024

A tribute to four programming languages

This post is a tribute to four programming languages: FORTRAN, PASCAL, C++, and LISP.

The algorithm that was programmed calculates the roots of the second degree polynomial:  

ax2+bx+c=0


The document containing the four codes: "programa raices.doc", was written in August, 2001, and it is available in GitHub .

Enjoy!


lunes, 8 de abril de 2024

Utilizing Ordinary Differential Calculus to Detect R-Waves in Standard EKGs

 


The electrocardiogram (ECG or EKG) is a vital tool for analyzing heart health. It captures the electrical activity of the heart, with distinct peaks and valleys corresponding to different stages of the heartbeat. 

Willem Einthoven invented the first practical electrocardiograph in 1895 and received the Nobel Prize in Physiology or Medicine in 1924 "for the discovery of the mechanism of the electrocardiogram".

The EKG signal typically consists of several waves: P-wave, QRS complex, and T-wave.

One crucial element of an EKG is the R-wave, which represents the peak of ventricular depolarization: the moment when the main pumping chambers of the heart contract. Detecting R-waves accurately is essential for interpreting cardiac rhythms and identifying abnormalities.

The R-wave typically occurs approximately 200 to 300 milliseconds after the onset of ventricular depolarization. Using the location of the local minima in the first derivative, we can estimate the time of arrival of the R-wave.

While there are numerous methods for R-wave detection, employing ordinary differential calculus can provide a robust analytical approach. 

Let's explore how ordinary differential calculus can help us pinpoint these R-waves.

DISCLAIMER

This approach was tested only in an academic environment, as part of a lecture about Electrocardiography. See Course below.

The Math Behind the Beat

An EKG recording can be thought as a function of time, with voltage on the y-axis and time on the x-axis. The R-wave corresponds to a peak in this voltage function. In Calculus terms, a peak signifies a maximum of the function.

The first derivative of the EKG function represents the rate of change of voltage over time. At the peak (R-wave), this rate of change will be zero. So, finding the maximums of the first derivative will lead us close to the R-waves.

EKG signals are often noisy. The first derivative might have minor fluctuations around the true maximum. To refine our detection, we can take the second derivative. The second derivative represents the rate of change of the rate of change (think acceleration).

At the R-wave, the voltage is at its maximum, so the first derivative is zero. But just before the peak, the voltage is rapidly increasing, resulting in a positive second derivative. Just after the peak, the voltage starts decreasing, leading to a negative second derivative.

Therefore, identifying points where the second derivative transitions from positive to negative will pinpoint the exact location of the R-wave with greater accuracy.

Putting it Together

By analyzing the EKG signal with these derivative properties in mind, we can identify R-waves:

  1. Find Local Maxima in First Derivative: Scan the first derivative for points where the value reaches a maximum. These points correspond to potential R-waves.
  2. Verify with Second Derivative Minimum: For each potential R-wave identified in step 1, check the corresponding point in the second derivative. A true R-wave will have a minimum value at that point in the second derivative.

Benefits and Limitations

This method offers a simple, calculus-based approach to R-wave detection. However, it's important to consider limitations:

  • Noise: Real EKG signals can be noisy due to muscle movement or electrical interference. These can introduce false maxima/minima in the derivatives, requiring additional filtering or noise reduction techniques.
  • ECG Variations: EKG morphology can vary between individuals, and some abnormal heart rhythms might not exhibit the classic R-wave characteristics. More sophisticated algorithms might be needed for robust R-wave detection in such cases.

Conclusion

Understanding R-waves is essential for EKG analysis. By utilizing the concepts of maxima and minima in ordinary differential calculus, we gain valuable insight into the electrical activity of the heart and its rhythm.

Ordinary differential calculus provides a valuable tool for understanding and analyzing EKG signals. By focusing on maxima in the first derivative and minima in the second derivative, we can pinpoint the crucial R-waves, offering valuable insights into heart function. 

However, it's crucial to acknowledge the limitations of this approach and consider real-world complexities for robust R-wave detection in medical applications.

Then, further research and validation are necessary to refine and optimize this method for real world applications.

You can find two versions of the source code: 1) MatLab/Octave, and 2) Python in my GitHub repo Signal/EKG/R-Wave 

Course lecture: 

Cursos/Electro_Medicina/elecmed05_2 Electrocardiografía.ppt

Note: This presentation is written in Spanish.

Recommended lectures

1.  Bronzino,J.D. (Editor)  “The Biomedical Engineering Handbook, 2nd Ed. IEEE Press, 2000 Chapter 13 “Principles of Electrocardiography”

2.  Carr,J.J y Brown,J.M. “Introduction to Biomedical Equipment Technology” Chapter 8 “Electrocardiography” pp 197-233

3.  Del Aguila, C. “Electromedicina” Ed. Hasa, 1994 Capítulo 8 “Bases de la Electrocardiografía” (pp 129-159) y Apéndice III Electrocardiógrafo (pp 475-499)

4. Webster, J.G. (Editor) “BioInstrumentation”, 2003



martes, 27 de marzo de 2018

Why Data Scientists prefer Python?

I usually write code in Java and C. Then, to start answering this question, I performed a simple test by writing the same code in the three programming languages: C, Java, and Python. My goal was to understand the differences among performance of of the three approaches, specifically regarding CPU and memory usage.

Then, I selected an algorithm that can be easily ported to Python too: Retrieving the second maximum value among three random positive integers. In order to get valid timing measures, this process was repeated 1000000 times. (+)

It must be pointed out that this is not a real benchmark test, because that wasn't what I was looking for. Even, the code does not make any input/output or network operations.

All tests were done in the same environment(!):
$ uname -a
Linux <this> 4.15.10-300.fc27.x86_64 #1 SMP Thu Mar 15 17:13:04 UTC 2018 x86_64 x86_64 x86_64 GNU/Linux 


The executables were created by using the following tools:
C:
$ gcc --version
gcc (GCC) 7.3.1 20180303 (Red Hat 7.3.1-5)

Java:
$ javac -version
javac 1.8.0_161
 
$ java -version
openjdk version "1.8.0_161"
OpenJDK Runtime Environment (build 1.8.0_161-b14)
OpenJDK 64-Bit Server VM (build 25.161-b14, mixed mode)

Python:
$ python --version
Python 2.7.14


The size of the codes, excluding JAVA Virtual Machine and Python Interpreter, are:
C: 8568 bytes, about 35 lines of code.
Java: 2721 bytes, about 40 lines of code.
Python: 1486 bytes, about 30 lines of code.

The commands to run the programs were:
C:
$ ./secondmax_c 1000000
Java:
$ java -jar SecondMax.jar 1000000
Python:
$ python secondmax_p.py 1000000

The results, in seconds, are:
C:
Best time: 0,038512 s
Worst time: 0,041267 s
Average time of 10 samples: 0,039393 s (*)
Java:
Best time: 0,190164966 s
Worst time: 0,213463167 s
Average time of 10 samples: 0,1982906219 s (*)
Python:
Best time: 14,326186 s
Worst time: 14,948839 s
Average time of 10 samples: 14,6719912 s (*)


So, if these tests show that the worst performance in C was around 347 times faster than the best performance in Python, and the worst performance in Java was around 67 times faster than the best performance in Python: Why do data scientists prefer Python? 

The answer is straightforward, and it might sound pretty obvious, even before testing: Python is easiest.

C and Java demand professional programming skills that must be acquired after a personal process of growth that sometimes may be slow.

However, a data scientist with her/his computer can understand and develop an acceptable Python code quickest.

Furthermore, Python has many libraries and solutions that can be freely included in the applications, reducing the development time, and therefore, the man hours getting a valuable answer to the different issues.

Anyway, I still have two open questions:
1.- What should be the best performance's option from System Administrator's point of view?
2.- Which one is the choice if the target environment of the application is an embedded system? In example: 454 MHz ARM 9 CPU with only 128 MB of RAM?

Notes:
(+) If you want testing my codes in your environment, just write a comment below and I'll send these. 
(!) To avoid comparing apples to oranges, I have omitted in this post the results after testing the three codes in different environments. It was tested too the performance of the C version in an embedded system: 454 MHz ARM 9 CPU with only 128 MB of RAM; and the performance of the PySpark version in a Cloudera 5.12.0-0 cluster.
(*) The average performance was only included because this is a standard measurement. I think that the value of the "worst time" is the most accurate predictor of the real time response capabilities of an executable program.


 

lunes, 30 de octubre de 2017

The Sentient Enterprise

I just have read the innovative book "The Sentient Enterprise", written by Oliver Ratzesberger and Mohanbir Sawhney.

This book is a "must read" for all of us involved in the improvement of the decision processes.

To clarify why, it is worthy to mention only two fundamental IT problems that are analyzed inside the book:
  • Why the enterprises are wasting human resources, time, efforts, their computing power, and money by unnecessary duplicating data?
  • Why the IT analytic solutions are usually delayed, and many times become obsolete when they are finally finished?
In order to help solving these issues, Ratzerberger and Sawhney introduce a model with five stages:
  1. The Agile Data Platform
  2. The Behavioral Data Platform
  3. The Collaborative Ideation Platform
  4. The Analytical Application Platform
  5. The Autonomous Decisioning Platform
This approach decomposes the huge problem: better decisions, and allows to develop specific solutions for each stage; thus making the whole process affordable.

I specifically recommend reading in detail chapter 7 "The Autonomous Decisioning Platform", and chapter 8 "Implementing Your Course to Sentience".

These chapters expose important guidelines, and introduce a perspective that matches most of the research topics behind the Autonomous Intelligent Systems (AIS)

lunes, 25 de septiembre de 2017

How useful are Recommendation Systems?

Typically, recommendation engines and systems enhance the user experience, because they assist us in finding information, reduce search and navigation time, and increase our satisfaction. 

However, I am still receiving recommendations about options to buy vacation packs, books, movies, music, etc.; that I had reviewed more than one year ago. Even worse, I receive friend suggestions because they are friends of someone that I know… Why? Really, most of the time I am not interested in receiving those kinds of recommendations… 

Therefore, I asked myself: How many people accept and follow these recommendations? 

Unfortunately, I don’t have access to all the required data in order to evaluate this; but according to Symeonidis and Zioupos (“Matrix and Tensor Factorization Techniques for Recommender Systems” ISBN 978-3-319-41356-3):
  • Amazon – 35% of product sales come from recommendations in Amazon.com  
  • Netflix – 66% of movies rented in Netflix.com are recommended  
  • Google – 38% more click-throughs are generated from recommendations in Google news
In my opinion, the main reason of this behavior can be traced back to the origin of the algorithms that have been used to create recommendations: clustering, ranking, scoring, pattern matching, etc. This means “the History”, usually understood as Big Data, OLAP, OLTP, etc. But it is clear that this "historical approach" demands many resources: storage, computing power, and time. 

But: What happen if we’re looking for recommendations about the outcome of a “random” process? 

The situation becomes harder if we don’t have enough information about the process itself. Let’s put it on an easy way: We need recommendations that could not be related to the previous history of the process. 

Example: A recommendation with probability 0.7 is not a winning one in gambling. Unfortunately, the number of alternatives experiences an exponential growth in order to achieve a greater probability. It is difficult to create such recommendations using "brute force" algorithms, and the task will demand the most powerful computers. Even worse: There is not heuristics to "prune" the decision tree.

This means that the inputs to the recommendation’s process are not well suited because the data could be either too poor or too much, the representation of the recommendation’s knowledge has not been identified as it should be, the inference rules could not be useful because there is not a previous experience about the behavior of the process, and the explanation about how the recommendations were created does not allow to identify the whole reasoning process: the domain is not well formalized. 

Then, I'm asking myself: How to predict the future behavior of a system whose previous history might not be relevant to the predictive process?

I believe that a new approach is needed to the recommendation’s process. This approach should redefine how to analyze the inputs, how to build new models of knowledge’s representation, how to propose different inference mechanisms that might not be well suited and computed by Turing’s Machines, many alternatives of solution, and explanations about how these recommendations were reasoned that may not match the common sense.

As a conclusion, I think that research in Artificial Intelligence should include more than machine learning, deep learning, and neural networks; because the focus of the problem should consider not only the data -the history-, but also the mechanisms to extract real knowledge of them: models, representations, inference rules, and explanations.

This path will lead us to “skilled intelligent systems”: solutions that can be transferred from one domain to other domains, and recommendations that will be really useful.   

viernes, 15 de septiembre de 2017

Comentario - Comment

Estimados lectores:

Cuando oficialmente dejé de trabajar como Profesor Titular en el año 2010, publiqué mis clases en Español para dominio público y sin fines de lucro en OneDrive < https://1drv.ms/f/s!ApPTAVJ07A-CgRQ2NbUjw6U1nDnB > .

A lo largo de varios años he visto dichos materiales referenciados y reproducidos por diversos sitios como SlideShare también sin fines de lucro. (Ejemplo en SlideShare: < https://www.slideshare.net/fuvylvp/almacenes-de-datos-olap-y-minera-de-datos?qid=2634f8ee-fb1d-4c57-aec0-a0add8ce77e6&v=&b=&from_search=1 > )

Sin embargo, hoy encontré un sitio que se declara "sin fines de lucro", pero que pide "apoyo al compartir y/o descargar" uno de mis documentos: "Aspectos Avanzados de la Tecnología de Objetos".

Personalmente considero que esta información debe ser conocida por todos.

Agradezco de antemano cualquier comentario ó sugerencia sobre acciones futuras al respecto.

Dear readers:

When I officially quit working as a Professor in 2010, I published my lectures in Spanish for public domain and non-profit use in OneDrive <https://1drv.ms/f/s!ApPTAVJ07A-CgRQ2NbUjw6U1nDnB >.

Over the years I have seen such materials referenced and reproduced by various non-profit sites such as SlideShare. (Example in SlideShare: <https://www.slideshare.net/fuvylvp/almacenes-de-datos-olap-y-minera-de-datos?qid=2634f8ee-fb1d-4c57-aec0-a0add8ce77e6&v=&b=&from_search= 1>)

However, today I found a site that is declared "non-profit", but "needs support to share and download" one of my documents: "Aspectos Avanzados de la Tecnología de Objetos".

Personally I consider that this information should be known by everyone.

I appreciate in advance any comments or suggestions on future actions in this regard.

Dr. Juan Jose Aranda Aboy